Stream and motion robustness: the shipper writes outside the lock, a
clamp inside an armed window faults, the X/Y soft limits follow the
home, the machine's settings are pinned, the homing keys are clamped
(grblHAL-glowforge 511ff25). Bench-proven with motion.soft-limits and
motion.deadman on the bench reference.
The GRBL driver's stream engine and its motion envelope change
(grblHAL-glowforge: the shipper writes outside the lock, a clamp inside
an armed window faults, the X/Y soft limits follow the home, the
machine's settings are pinned, the homing keys are clamped). This commit
carries the host rules and the catalog tests that hold them; the driver
commit follows, because its CI fetches these harnesses unpinned.
scripts/bench/laser_stream_test.py:
- Rule 28: a 300 ms producer stall while armed faults the stream with
ALARM:17, the kernel sees no step burst (at most the planned steps per
100 ticks), the stream ends dark, the latch sideband ends on the lock.
The same stall unarmed is a warning: the move completes with every
step, the clamp visible as the burst the kernel counts.
- Rule 29: a 300 ms stall of the sink's write leaves the producer on
pace: no clamp, every step, lit through, dark at the end.
- The stand-in engine takes GFSINK_STALL_MS and GFSINK_WRITE_STALL_MS,
null-sink only.
scripts/bench/z_envelope_test.py:
- Rule 10: homed (a gfcloud home), a program move past X max, Y max or
the near edge alarms with ALARM:2 before any motion, a jog past the bed
is refused with error 15, a move inside the bed runs, and a $20 write
keeps the limits. The core repeats the last error for the line after a
refused jog until an empty line clears it, so the rule sends one.
forgetest/forgetest/suite/motion.py:
- motion.soft-limits (kind auto, no emission): homes through the cloud
suite's gfhome homing when the machine is not homed, then the three
refusals (ALARM:2, the kernel counters still), the refused jog, the
inside move, and the return to the corner read at rest.
- motion.deadman phase 2b: a 100 ms SIGSTOP mid-move, inside the
kernel's queue: no underrun, the controller's log warns of the clamped
late events, the move completes with every step (read at rest), the
latch stays locked. The armed clamp is proven on the host (rule 28).
forgetest/forgetest/suite/cloud.py:
- gfhome_homing drains the driver's answer to $H once the session ends:
it sits behind the status reports and passed for the reply to the
caller's next command (a setting read as None).
forgetest/forgetest/baseline.py:
- The hand-back reads the position counters at the kernel's own
microstep mode (cnc/x_mode, read before the sysfs restore puts the
settings' mode back). At the x8 constant, an x32 machine's 30 mm read
as 120 mm, beyond the return bound, and the displaced head was left in
place. The dead band scales the same way. tests/test_baseline.py holds
both.
Proof. Host: stream rules 1 to 29 and z_envelope rules 1 to 10 against
the null-sink driver, the forgetest unit tests. Bench reference:
motion.soft-limits passed (X 495 and Y 279 refused with ALARM:2 and
0.000 mm of kernel motion, X-1 the same, the jog error:15, the inside
move ran, back at the corner 0.0/0.0 mm); motion.deadman passed with
the new phase (92 late events clamped, max behind 92.3 ms, no underrun,
30.0 mm counted, latch locked) and the hand-back jogged the head back
under the x32 scale.
The emission gates: the per-tick fire gate, the latch with an owner,
the cooling verdict's two tiers, and the arm-flow gates
(grblHAL-glowforge 16d2b9e). Bench-proven with laser.verdict-cut and
laser.armed-kill on the bench reference.
The GRBL driver's emission gates change (grblHAL-glowforge: the per-tick
fire gate, the latch with an owner, the cooling verdict's two tiers, and
the arm-flow gates). This commit carries the host rules and the catalog
tests that hold them; the driver commit follows, because its CI fetches
these harnesses unpinned.
scripts/bench/laser_stream_test.py:
- Rule 24 is the verdict's pause tier: the client holds the job under the
open window, the first deceleration runs lit to the stop (the dark lead
before the stop is at most 1000 ticks: the producer's lead plus one
shipper period), a resume under the standing verdict moves dark and is
held again, the clean verdict resumes lit with no press, and the latch
sideband carries only the arm's unlock and the program end's lock.
- Rule 25 is the fail tier: AIRFLOW mid-cut ends the job with ALARM:3,
the stream ends dark well short of the line, the sideband ends on the
lock, and a resume under the clean verdict that follows resumes nothing.
- Rule 26: a sender change mid-M3 holds the job with the deceleration
dark: the gate follows the window on every tick.
- Rule 27: a verdict that goes stale holds the job at the cache's own
expiry, lit to the stop, never a poll later; the engine's return
resumes lit.
- The stand-in engine publishes the verdict name and has a stale mode;
the session steps gain expect_text, reconnect and a wait_state timeout;
every session reads the latch sideband (GFSINK_LATCH_LOG).
scripts/bench/laser_lifecycle_test.py:
- Rules 11 to 15: the pause tier resumes with no press and no prompt, the
fail tier ends the job and nothing resumes it, a sender change during a
re-arm cancels it, a jog does not hold the window open, and a press
counts only after the button has been seen up. The stand-in engine
takes a live verdict dict. start_armed_move waits for a fresh prompt
and a fresh armed message: a press that lands before the wait has
begun is not consent, and the old stale match let one land early.
forgetest/forgetest/suite/laser.py:
- laser.verdict-cut (kind live, one press): a 40 mm M3 line; the test
pauses the daemon for 3.5 s so the verdict goes stale (the settings
route is idle-gated and the engine reloads its gates at a session
start, so no setting can trip a pause mid-cut; the crash tiers need a
physical knock). The controller must hold with the SoC latch and the
hardware button latch both clear in every sample, emission must read 0
before the resume, and the clean verdict must resume the cut lit with
no press and no prompt; M2 disarms as usual.
- laser.armed-kill asserts that the respawned controller comes up with
the latch still locked.
Proof: both harnesses pass against the driver change on the host, and
the forgetest unit tests pass. On the bench reference, laser.verdict-cut
passed with no gap in the cut (held at +1.71 s after the pause began,
both latches clear in every sample, emission 0 at +2.06 s after the
hold, resumed lit at +4.0 s with the beam detector 678 counts over idle,
kernel drift 0.0 mm, disarmed 0.1 s after Idle) and laser.armed-kill
passed (emission 0 at +1.8 s after the supervisor's stop and +2.1 s
after the SIGKILL, latch locked, respawned with the latch locked, button
dark). Two earlier verdict-cut runs shaped the driver: a pause that
locked the latch resumed dark, because a lock sets the hardware button
latch, and a hold taken a poll after the gate closed left a several-mm
gap.
Catalog: one test added and one extended; both cover src/** of the
driver through the existing laser covers.
.gitattributes sets text=auto with eol=lf, so every text file is
stored and checked out with LF, and a patch keeps its bytes. The files
that carried CRLF from a Windows editor are renormalized. No content
changes.
The commission.* tests are setup.* (subsystem "setup"): suite/setup.py,
setup_dark.py, setup_sheet.py, and the host tests test_setup_*.py. The
coverage maps name src/setup.* in place of src/commission.*, the record
is setup.json, and the bench seed in forgetest.init creates
/run/forgefirm/setup-override.
setup.check-sensors follows the check as it is now: no question is
asked and no setting is written. The test reads the settings before
and after and fails on any change, and fails at once if the check
opens a prompt. The host test drives the fake daemon with no prompt
and proves both outcomes.
Proof: test_setup_dark.py, test_setup_sheet.py, and
test_setup_suite.py pass (46 tests). The coverage lint names
src/commission.c and .h as uncovered until the forgectrl pin moves to
the revision that carries the rename.
FORGEFIRM_RELEASE 0.0.4. The acceptance record that authorizes this
build: campaign c-20260912182427-e1ee on dev image 20260912180956
(manifest 8bb5b6f3), 85 tests, 85 satisfied, 0 required, exported
2026-09-12T19:21:35Z. The release notes in releases/v0.0.4/notes.md
name this the first public beta; the pipeline takes a release's
notes.md, when it carries one, in place of GitHub's generated notes.
forgectrl 0.1.24 = 2380e07: every fan duty write is read back and
retried, a duty a device lost is put back by the tick, and the airflow
fault names the duty commanded and the duty in force. The corner card
of the sheet on the bench reference was held as a slow air-assist fan
after a run-duty write the head never took; nothing checked.
Acceptance: cooling.fan-duty-readback (auto, grbl) opens an M8 session,
reads every fan's run duty back, writes the head's air-assist register
back to the idle duty and the exhaust PWM to zero behind the engine's
back, and holds both to their run duties again within a few ticks, each
loss named in the log, the session OK to its end, the idle duties after
M9. Covers the head driver too. On the bench reference with forgectrl
0.1.24 hot-deployed it PASSED: both duties back after 0.5 s.
forgectrl 0.1.23 = cb8c140: the dev-server mirror test holds the mock's
release replies to the keys reply_release packs; the daemon is the one
0.1.22 carried. CI is green on this revision.
forgectrl 0.1.22 = 0235a88: the release check reads the GitHub releases
API and never requests the firmware file's URL, the daemon checks daily,
and the panel raises a per-release dismissable alert and runs the install
from one dialog.
Acceptance: update.release-check (auto) exercises GET /update/release,
POST /update/check (a machine with no route to the API answers 502, which
the drill records and steps over), the v<semver> shape and `new` of a
published release, and the dismissal round trip, and puts the dismissal
back. forgectrl.auth's unauthenticated-write list gains /update/check and
/update/dismiss. Coverage lint: 84 tests, 0 uncovered paths; the
forgetest unit tests pass (373). On the bench reference (dev image
20260911203113 with forgectrl 0.1.22 hot-deployed) every check of both
tests passed; the runs were marked FAIL only by the hand-back baseline,
because the controller is gated until the changed privacy advisory is
accepted again.
Pin forgectrl at 92cead6 (0.1.21). The published-release check reads
the release tag from the first redirect hop instead of the end of the
chain, where the asset store's URL carries none; every image through
v0.0.3 reported "release server error (HTTP 200)" against a published
release. The forgectrl commit carries the proof (relcheck_test in its
CI). The update.slots-and-signature covers map names the new
src/relcheck.c and src/relcheck.h so the coverage lint stays whole. A
catalog test of the check against a published release is held for a
later change.
FORGEFIRM_RELEASE takes 0.0.3. The file sits outside the layer content
hash, so the bump changes the version and invalidates no acceptance
result.
releases/v0.0.3 carries the artifact the bench exported for this image:
campaign c-20260911182256-078f on 20260911172215 (dev), manifest identity
fb24c3f4508ae54592f8d0b98ed0b2a07a933274980d8c1e4e5a967618ef9bd8, 83
tests, 83 satisfied (57 inherited), 0 required, release authorized. The
release gate recomputes every catalog test's domain fingerprint from the
manifest inside the release rootfs and signs only when the recorded
results agree.
The release carries, since v0.0.2: the lens frame in one place, so every
Z a commissioning card sends comes from the settings the controller
opens its Z limit from (the tail at the head's reference, the focus
window written before the card's controller starts, every program judged
against the reach before its first line goes out); the commissioning
sheet acceptance test run as a fresh machine's; and the CPU percent on
/status that holds across a status read inside the same scheduler tick
as another reader's. Components: forgectrl 0.1.20, grblhal-glowforge
0.1.11, forgefirm-app 0.1.28+git, kernel-module 0.0.5, meta-openglow
ced2af2.
forgectrl 864b8da repeats the last CPU percent when a /status read finds
the /proc/stat counters unmoved, the race that failed
forgectrl.panel-serves on the release candidate when the test's second
read shared a scheduler tick with a panel poll. Every test covering
forgectrl re-runs.
forgectrl 3e54612 renames the lens_test step in build.yml, whose name
held a colon that YAML read as a mapping, so CI on f8ddb17 ran nothing.
The source is f8ddb17's; the pin moves so the release names a commit
with a green build.
forgectrl f8ddb17 puts the lens frame in one place and writes the focus
window before the card's controller starts, after two commissioning
cards ended in ALARM:2 on a Z the wizard sent from one source while the
controller's Z limit stood on another. The acceptance run had passed
only because the bench's settings already held the stops from an
earlier focus run, so the run is now a fresh machine's.
commission.sheet clears the three lens settings inside its Restore
before the cards, checks that the frame runs in the fallback window,
and, after the focus card, that the settings hold the window the ladder
ran in (the stops found or the fallback), that every ladder height lies
in that window's reach, and that the program served now agrees with
/status. Every served program's Z is checked against the reach /status
reports before the card starts, so a stray Z fails the test with nothing
burned. The test covers src/lens.*. The host mock in
tests/test_commission_sheet.py mirrors the daemon (the /status lens block
from the settings, the ladder served from the settings, the window
written at the focus start); two regression tests reproduce the defects:
a focus result naming a window the settings do not hold, and a served
program with a Z beyond the reach. 10/10 green.
scripts/bench/z_envelope_test.py, the grblHAL CI harness, adds the
referenced-lens cases: with forgectrl's marker and the shared settings,
the fallback window and a 14/20 window run to the ends of their reach
and two half-steps past either end alarms, and a count of 41 falls back
on its side alone. Passed on the null-sink build. The bench page's
description of the harness follows.
The forgectrl pin moves to f8ddb17 (0.1.18); every test covering
forgectrl re-runs.
FORGEFIRM_RELEASE takes 0.0.2. The file sits outside the layer content
hash, so the bump changes the version and invalidates no acceptance
result.
releases/v0.0.2 carries the artifact the bench exported for this image:
campaign c-20260910181641-11a2 on 20260910181308 (dev), manifest identity
a31c820d26d09155d2bf7c629389c5dba5e55d2b52dbbc0ccab5440b187b3b3c, 83
tests, 83 satisfied (57 inherited), 0 required, release authorized. The
release gate recomputes every catalog test's domain fingerprint from the
manifest inside the release rootfs and signs only when the recorded
results agree.
The release carries, since v0.0.1: every machine named after its own MAC
with mDNS dropped, the fan run posture a diagnostic measures, the
hand-back rules the bench now enforces, the flow-load tail ending at the
coolant peak, and the stream flag a dead session leaves behind cleared by
the session that takes the device. Components: forgectrl 0.1.17,
grblhal-glowforge 0.1.11, forgefirm-app 0.1.28+git, kernel-module
0.0.5, meta-openglow ced2af2.
grblHAL-glowforge 0.1.11 (1f2ea9c): a controller taking the pulse device
clears a dead session's cnc/streaming flag, the way it already acks that
session's stale underrun. laser.armed-kill kills the controller mid-fire,
and the machine came back with the kernel still believing the killed
session was feeding it.
laser.disarm-in-hold streams a 40 mm move, feed-holds it two seconds in,
waits out the disarm grace, and then resets out of the hold. The recovery
put the laser and the controller back and left the head where the hold
had caught it: 11.34 mm along on the bench reference, 605 counts, which
the hand-back jogged out and reported.
The head goes back now, by the distance it actually travelled rather than
the distance the move asked for: the hold catches it at a slightly
different point every run - ten millimetres nominal at F300 for two
seconds, 11.34 measured - so the kernel counters before the move and
after the reset are what the return jog is built from. Where it stopped
and where it ended go into the evidence.
This is not the fault the M5 rapid job had. That job's moves netted plus
twenty by construction; this one stops part-way on purpose, and the
recovery simply never returned it.
laser.armed-kill has the same shape, a kill mid-fire that stops the head
where it stops, and is left alone until a run says whether it needs the
same treatment.
laser.m5-rapid-dark cut 20 mm out and then ran the two rapids the test is
about, one back and one out again. The three net to plus 20 mm, so the
job ended with the head 20 mm from where it started, every run: 1067
counts at 53.333 per mm, which the hand-back jogged out and reported.
That one is dirt, and the position dead band was right to leave it alone
- it is 20 mm, not the step a return rounds to.
A third rapid, back 20 mm with a short dwell, ends the job where it began.
It sits inside the sampling window and after the M5, so it is one more
rapid that must be dark, which is what the test already asks of the other
two: the assertion is wider, not narrower. The operator's clearance step
is unchanged, because the head still needs its 20 mm of +X.
Every job list in the suite that moves in G91 now nets zero on X
(cooling.py, laser.py twice).
cloud.verdict-hold printed to completion and then failed its hand-back on
head/z_mode=0. That job's header carried ZSmd 0, gfhardware's
set_mode_from_puls turns that into z_mode 0, and z_mode 0 is full step
(lenshome.c): the cloud client had set the lens the way the pulse file it
was playing asked. Nothing was left behind - the machine did what the job
said.
The baseline already skips nine attributes in cloud mode, for the reason
written above the list: the cloud client sets its own values for them
from every pulse header, and forcing the GRBL values under it would be
the baseline configuring another controller's machine. head/z_current and
head/z_mode are the same thing and were not on the list. They are now, in
cloud mode only; in GRBL mode the pair is still checked against 1 and 1,
which is the half of this an exemption could quietly break, so a test
holds it.
Host-proven: 89 baseline unit tests.
A leftover is dirt a run left behind. Four of the checks were reporting
something else: the machine part-way through work of its own, which it
finishes on its own a few seconds later. Each one failed a run that had
done nothing wrong.
laser.arm-wait-lid found it. The test passed every check it makes and
failed on cool=smoke/armed=False/hold=False. The smoke clear is the
engine's post-job work - run duty for smoke_s after an armed session,
timed, ending by itself (cool.c, Cool_Smoke) - and the machine was idle
fifteen seconds later without being asked. Armed and hold were both
false: nothing had been left.
The four, and what each was catching:
cool a phase alone (the smoke clear, a cooldown). Armed or
holding stays a leftover: the run left a job alive and the
engine is keeping the fans up for it.
controller the supervisor's own start. A takeover ends by starting
forgectrl again, and the respawn runs the liveness probe
and the lens reference before it reports running and
verified. The run had put it back.
state the ring draining to the end of a job. An underrun stays a
leftover and is still acknowledged with cnc/stop: that one
is the run's.
leds read_led read brightness, write_led writes target, and the
smooth trigger fades brightness toward target. An LED the
machine had already released still read lit mid-fade, and
the restore called itself done before the fade had moved.
Judged on target now: a run that left the button lit left
a target standing.
The waiting, the restoring and the reporting of anything a run actually
left are unchanged. What changed is which readings count as dirt.
Host-proven: 87 baseline unit tests, two of them new - a fading button
LED is not a leftover, a lit one is.
The airflow check reads the purge fan's off current with the fan off, and
the stand-down that follows commands it back on. The guard that proves
the machine was handed back whole read the draw immediately, so what it
got back was the off current the check had just measured: on the bench
reference, 74 against a 300 floor, with the fan drawing 631 a moment
later. The check failed for having worked.
The current follows the command; it does not arrive with it. The guard
now waits for the draw to reach the floor, up to PURGE_SPINUP_S, the way
the controller below it is already waited for, and logs how long it took.
It is no weaker: a fan that never reaches its floor still fails the
check, and the message now says how long it was given.
Pin: forgectrl 0.1.17 (a9e7073, the air assist and the purge in the
diagnostic run posture - the reason that check measured an idle fan and
wrote a floor from it).
The hand-back is the promise that a run leaves the machine where it found
it. Two of its checks reported instead of restoring, and the machine sat
in the state the run left it in.
The cooling engine: a run that ends without ending its job leaves one
alive behind it - the engine armed and holding for a job that is never
coming back, the fans at run duty. The check waited two minutes for that
to resolve itself, which it cannot, and wrote "failed: still
run/armed=True/hold=True". On the bench reference the fans then ran for
an hour.
The pulse ring: bytes the last job never played sit there, and the next
run replays them before its own. The check refused the return jog and
wrote "clear the ring (controller restart) before moving" - the
instruction, to a log, instead of the action.
Both end the same way: stop the controller and let the supervisor bring
it back. The job goes, the arm and the hold go with it, and the ring is
empty on the way in. stand_down() does that and proves it settled;
the cooling check calls it when the engine will not idle on its own, and
the return jog calls it when the ring has residue, refusing only if bytes
survive a restart. A leftover still reports what it found - the record of
what the run did is the point - but it reports having fixed it.
Host-proven: 62 baseline unit tests, including one that the hand-back
calls the stand-down for ring residue rather than describing it.
Run.snapshot() computed elapsed_s from the current time on every call,
whether the run had ended or not. The page shows the last run until the
next one starts, and it polls, so a finished run's figure went on
counting: the result badge said FAIL beside a number still climbing, and
the run read as still going. On the bench reference a test that ended
after 5082 s was showing 5242 s and rising.
The figure a finished run should carry was already recorded next to it:
finished["duration_s"], fixed when the result was written. snapshot()
now returns that once the run has ended, and the live count only while
it is running.
Host-proven: two unit tests - a finished run's clock reads its duration
and does not move across a poll, a running one's still climbs.
An empty form field never reaches the request, so POST /settings with an
empty value in the body carries no key at all and the daemon answers 400,
"no known setting in request". Only the query-string form gets an empty
value through. Measured on the bench reference:
POST /settings -d "lid_policy=" -> 400, the value unchanged
POST /settings?lid_policy= -> 200, the key cleared
Three places in the suite already knew this and say so in a comment; two
did not.
motion.lid-policy-hold restored with the form body, so on a machine whose
lid_policy has never been set the restore was refused and the test failed
its own check. It now uses the query-string form for the empty value, as
the others do. This is the same test as the commit before: that one
stopped it from writing the word "cancel" where the machine had nothing,
and this one lets it write the nothing.
cloud.mode-switch had the older shape of the same fault: it restored
homing_mode to the literal "none" where the machine had it unset. On the
bench reference the restore branch never ran, because that machine homes
through the cloud, so it has not fired yet; on a machine with homing_mode
unset it would have. It now restores exactly what it found.
The rest of the suite was audited for this and is clean: camera.key-read,
cooling.critical-tier, cooling.tec-drive, cooling.crash-watch-plumbing,
cooling.fire-watch-tiers, cooling.floor-and-warm-up, cloud.verdict-hold,
forgectrl.settings, logs.levels and motion.xy-microsteps all either guard
the empty value or go through the shared Restore helper, which has always
done it correctly. commission.cloud-header-capture uses that helper too.
motion.lid-policy-hold read the setting as
was = (fc.settings() or {}).get("lid_policy") or "cancel"
and wrote `was` back at the end. On a machine that has never set the
policy the setting reads as the empty string and behaves as cancel, so
the `or` turned "unset" into the word and the test handed the machine
back carrying a setting it did not arrive with. The hand-back reported
it, restored it, and failed the test.
The value is now captured exactly, empty included, and restored as
captured; the default belongs to reading the value, never to writing it
back. The read-back check gets the same default, so it no longer compares
None with the empty string. lid_policy_in_force carries the effective
policy into the evidence, which is what the old expression was reaching
for.
Found on the bench reference, on the run after the position dead band let
motion.lid-cancel-home through. The four other places in the suite that
read a setting with `or` are safe: two default to the empty string, which
is what unset is, and two feed a check rather than a restore. The shared
Restore helper captures raw values and writes the empty string back for
unset, as this now does.
Only this test's earlier passes are invalidated: the change is inside its
own body, and the per-test source hash of the other eleven motion tests
is unchanged (checked against the file before the edit).
The post-run pass compared the kernel position counters exactly, so a
test that put the head back within a hundredth of a millimeter failed
whenever that distance did not round to the same step.
motion.lid-cancel-home found it on the bench reference. The test cancels
a job with the lid twice and the controller returns the head to the job
start each time, landing 0.038 and 0.037 mm out - the same figures as the
run before it, which passed. This time the two returns left four steps on
X, 0.075 mm at 53.333 steps per mm, and the hand-back called it a
leftover. Twenty-three runs of that test, all passing, on returns of the
same accuracy: whether the residue rounds to zero is chance, not a
property of the machine or of the test.
A difference inside POSITION_DEADBAND_MM (0.1 mm, five steps at x8) on X
and Y is now the quantization rather than a leftover. The head is still
put back, so nothing accumulates over a campaign - only the failure goes.
Z stays exact: the return never moves the lens, so a Z difference is
still a leftover and still unrestorable.
Host-proven: four new unit tests on the boundary (four steps and a moved
Z on either side of it, and an unreadable reading), and the forgetest
suite.
The sheet handed the bench actuator every press after the operator's
presence press, and the actuator pressed every time - into nothing. On
the bench reference all six cards were pressed before the card asked:
the frame by 8 s, the focus by 33, the floor by 10, the dose curve by 8,
the corner by 9, and the flow-load card by 69. The operator then pressed
all six himself, which is the opposite of what the ready gate promises.
arm_press() waits for hw.button_lit(), which is true when any button LED
is on, and burn() started that wait before run_check had even started the
wizard. The button is lit through parts of a card that are not the arm -
the lens reference, the program on its way to the controller - so the
wait ended at once, the press landed before the job waited for it, and
the thread was gone by the time the real cue came. forgectrl uses the
same predicate but only inside the job's own sample callback, with the
tube still dark, where a lit button does mean the arm.
The machine already says when it wants the press: a live check opens a
`press` wait prompt at that moment, and run_check sees every prompt. It
now presses there, through a new Ctx.press_now() - no LED read, no
waiting thread, no timing guess. That retires the per-card lit-timeout
column of CARDS, which existed only to give the flow-load card's coolant
settle enough room for a wait that was reading the wrong thing.
The four grbl-driven arm_press() callers in laser.py and cooling.py are
left as they are: they call it with the machine idle and its LEDs dark,
so the level read is the edge they mean. The same shape would bite them
if that ever stopped being true.
Pin: forgectrl 0.1.15 (6e69479, the flow-load tail ending at the coolant
peak and the way out of a finished setup page).
Host-proven: 357 forgetest unit tests, including two new ones - the
actuator presses on the prompt with button_lit stubbed false throughout,
so the LED is provably not consulted, and the press falls to the operator
without a takeover.
One name for every machine was wrong: an operator with two of them on a
network had one forgefirm.local, and mDNS does not work on many networks
at all. The machine now calls itself forgefirm-<xxxx>, from the last four
hex digits of its WiFi MAC address, and sends that name with its DHCP
request, so a network with dynamic DNS publishes it and a router lists
the machine by name. The name is the same at every boot, two machines
take different names, and no serial number leaves the machine.
forgefirm-hostname (new): reads the wlan0 MAC address (eth0 on a machine
with no WiFi) at S38 in rcS, after udev has probed the network drivers
and before poky's hostname.sh reads the file and before the network
starts. The rootfs is read-only, so the name is written through a
bind-mounted copy under /run/forgefirm. A bounded wait covers a slow
probe. hostname:pn-base-files is "forgefirm": the name before S38, and
the fallback when no MAC address can be read.
avahi is deleted - the bbappend, the daemon configuration, the service
file, the image install and the distro block. The address is the way in
that works on every network, and the DHCP name covers the rest.
forgefirm-banner: the marker lines are gone. "# ForgeFIRM addresses" and
"# end" delimited the address block inside /etc/issue, and getty prints
every line of that file, so both markers were on the console. The script
now keeps the image's own text in a second copy under /run/forgefirm,
captured once per boot before the first write, and renders the whole
banner from it. The block is the addresses alone: no mDNS name.
forgefirm-image.bb: the ForgeFIRM mark, under the OpenGlow one the base
image carries, with the version on the mark's own last line,
right-justified to the mark's last column. The mark is written once and
rendered per reader, because /etc/issue is parsed by busybox getty (a
backslash or a percent sign starts an escape, so the art goes in with
every backslash doubled) while /etc/motd is written out as it is. Widths
are measured in columns, not bytes: the color sequences take no room on
the screen. /etc/issue.net stays unused - the machine tells a client that
has not logged in nothing.
Acceptance: commission.mdns-announce is replaced by
commission.machine-name, which checks the name against the MAC address,
the bind-mounted /etc/hostname, the DHCP client's hostname option, the
banner's addresses, and that no mDNS responder is on the image; it covers
nothing by design, like the test it replaces. forgectrl.auth gains the
own-name Host check and its refusal with a domain on it. image.health
checks the /etc/hostname mount and the version on the mark's last line in
both files. commission.ssh-until-reboot asserts there is no
pre-authentication banner. commission_dark's lens coverage widens to
src/lenshome.* so src/lenshome.h is covered; the lint is clean at 83
tests.
Pins: forgectrl 0.1.14 (9e5330f, the hostname certificate and the Host
rule), meta-openglow ced2af2 (the DHCP hostname option and the motd mark)
in the kas lock.
Proven on the bench reference, hot-deployed and rebooted (image
20260910000208 dev): hostname forgefirm-b00a from MAC 2c:6b:7d:0d:b0:0a,
live and in the bind-mounted file; the DHCP client running with
-x hostname:forgefirm-b00a; the console banner and the motd carrying both
marks with the version aligned to the mark's last column, no marker line
and no .local name; forgectrl regenerating its certificate for the new
name. Host tests: 357 forgetest unit tests, forgectrl clean under
-Werror, tls_test and sanitize_test.
The acceptance page shares theme.css with the panel byte for byte, and forgectrl's copy gained the setup header's Download logs button. CI caught the drift, which is what that check is for; the pinned revision decides, so the copy follows.
The baseline has always examined the machine after every run and recorded
what the run left behind. It did nothing else with it: the leftovers went
to the log and the evidence, and the test still reported PASS. So a check
could measure correctly, walk away with the machine in a state nobody
chose, and be recorded green.
That is how the purge fan came to be left off by the airflow check. The
leftover was not even watched, but had it been, it would have been noted
and the test would have passed anyway, and an operator would still have
met the airflow hold at their first fire.
A post-run leftover now fails the run. One the baseline put back fails it
too: the restore is the bench cleaning up after a defect, not the defect's
absence. The message names what was left.
The baseline watches the head as well as the motion side now: purge air
on, which is how the machine idles, and the lens motor at its hold current
in half step, which the lens checks and the sheet cards take and must hand
back. The airflow check proves the machine is whole rather than merely
measured: afterward the purge fan must read commanded-on and must draw
above the floor the check itself just wrote.
The pin takes forgectrl 0.1.13 (a2d73ef), which restores the idle posture
after a diagnostic, fixes the lens session's takeover flag, and gives the
setup a Download logs button, since the panel's Logs tab is unreachable
until the setup is complete.
Expect this to find things. A test that has been handing the machine back
imperfectly has been passing until now, and the first campaign under the
rule is where that shows.
The guard keeps scripts/manifest-from-tree.py and forgefirm-image-manifest.bbclass computing the same identity, so it pins the skip list by value and the new suffix failed it. It now expects the version file in both places, and a second case proves the point the change exists for: setting the release number leaves the layer hash where it was. My miss: I proved the identity by hand and pushed without running the unit suite first.
Setting the release number was a platform change. FORGEFIRM_RELEASE sat
in forgefirm-image.bb, the recipe hashes as content of meta-forgefirm,
and a change to the content of a layer invalidates every acceptance
result. So a version bump threw away the campaign that was meant to
authorize that very release, and the number therefore had to be decided
before the image the campaign ran on. Nothing said so: the release-flow
page went straight from the kas configuration to the artifact and the
pipeline, while the gate quietly required the recipe value, the rootfs
stamp, the archive's meta-version and the tag to agree. v0.0.1 was cut
on a tree whose number happened to be right; the next one would have
cost a second campaign to discover the rule.
The number moves to forgefirm-release.inc, which carries it and nothing
else, and the manifest leaves that file out of the layer content hash
exactly as it leaves out the component pin files
(FORGEFIRM_MANIFEST_VERSION_SUFFIX, and the same list in
scripts/manifest-from-tree.py, which computes the identity on a
workstation and must agree byte for byte). release.sh reads the number
from the new file.
The version is metadata, not platform content, and this only makes the
manifest say what it already meant: the version string was already
outside the identity hash, and it was the file carrying it that defeated
that. Nothing is weakened. release.sh still requires the number to equal
the rootfs stamp, the .fw meta-version and the release tag, and
image.health still compares the stamp on the running machine with the
manifest's.
Proven: the tree manifest is byte-identical across a bump from 0.0.1 to
0.0.2 (identity a64e51b8e5ecca0af683d4f0 either way, the meta-forgefirm
layer hash unchanged), where before the two differed. bitbake resolves
FORGEFIRM_RELEASE=0.0.1 and FORGEFIRM_VERSION_STRING=v0.0.1 for the
release image through the new require, and the dev image still overrides
the string with its build timestamp.
Campaign c-20260909204732-3fa9 on image 20260909193551, 83 tests, 83 satisfied (62 inherited), release authorized. It replaces the artifact of the release that was withdrawn: that one authorized an earlier rootfs, and the gate compares against the rootfs it is asked to sign.
ForgeFIRM's own files live under /data/forgefirm; two configuration
files did not. The machine settings sat at /data/forgefirm.conf, in the
root of /data beside the factory's own files, and the cloud-mode
configuration sat at /data/etc/gfhome.conf, inside a directory the
factory owns. Both move:
/data/forgefirm.conf -> /data/forgefirm/forgefirm.conf
/data/etc/gfhome.conf -> /data/forgefirm/gfhome.conf
There is no migration: only the bench has ever run this firmware.
/data/etc now holds only the factory's wpa_supplicant.conf.
The acceptance check of the file modes reads the settings file at its
new path, and the two bench tools that read it directly follow. The
pins move to the revisions that carry the change, forgectrl also
bringing the fix that reads the module's disabled state as idle:
forgectrl 468ee21 (0.1.12)
grblhal-glowforge 9ee624b (0.1.10)
forgefirm-app 56f134a (0.1.28+git)
The lock moves meta-openglow to b7ad6d9, which pins python3-gfhardware
on the same revision. The four upstream layers stay where they were:
`kas lock --update` moves every floating repository, and a release is
not the place to take poky, meta-openembedded and meta-freescale along
for the ride.
A machine out of the box sat in the module's power-on state, disabled, for the whole of its first run, and every idle gate read that as busy: the setup's sensors check refused to start, settings writes answered 409, and the cooling engine held cooldown airflow from boot. forgectrl now reads disabled as idle (the state means no program in progress), with fault and underrun still busy until acknowledged and an unreadable state still failing closed.
The yocto-cold-build workflow and its kas/ci.yml overlay built the release image on a hosted runner as a reproducibility probe. It never ran to completion, its first dispatch (2026-09-09) stopped on the runner's user-namespace rule, and a probe nobody runs is a trap. Every Yocto build, the release included, runs on the build host; the release proof is the local pipeline (release.sh) and the bench campaign. The pre-publish checklist loses its self-containment line to match.
BitBake isolates the network of its tasks with a user namespace, and the ubuntu-24.04 hosted runner's AppArmor profile refuses that to an unprivileged process, so the cold build stopped before its first task (run 34381825302). The workflow lifts the restriction for the run; nothing in the layers or the image changes.
The kas configuration takes the pinned-remote meta-openglow block, with
its commit in the lock file (d655e1e, the read-only rootfs), so a fresh
clone builds the release without a sibling checkout. The lock keeps the
upstream layers where they were.
releases/v0.0.1 carries the acceptance artifact the bench exported for
this image: campaign c-20260909160235-7649 on 20260909150456, 83 tests,
83 satisfied, none inherited, release authorized. The release gate
recomputes every test's fingerprint from the manifest inside the release
rootfs and signs only when the recorded results agree.
forgectrl now holds the motion check while a lid or the interlock is
open instead of starting the controller unverified: GET /mode reports
controller "waiting" with why, the button blinks amber, and the check
runs when the enclosure closes. motion.gate-waits-for-lid drives the
fixture's lid channel: the lid opens, forgectrl restarts, /mode must
read waiting with why naming the lid, no pid, motion unverified, the
button amber (sampled over a blink period: the smooth trigger's target
reads 0 through the off half) and no probe line in the log; the lid
closes, and the controller must come up verified with MOTION OK on the
first probe.
Proven on the bench reference: PASS, the controller verified 6.5 s
after the lid closed.
The rootfs mounted read-write, so a slot ran with its own files open to
change, and the factory-slot mounts rode along on the release image.
Both images now carry the read-only-rootfs feature: the ro root line and
the rcS default, the volatile links made at rootfs time, a writable copy
of /var/lib at boot, a build failure for a post-install that needs the
machine, and the removal of shadow, base-passwd, update-rc.d and
update-alternatives.
What must last or change at run time is handled file by file:
- forgefirm-users renders the four account files from the record into
/run/forgefirm/accounts and bind-mounts each copy over its /etc file
(useradd and the rest are gone with shadow); a render writes through
the mount, and the image's own files apply until the first render.
- forgefirm-banner bind-mounts a copy of /etc/issue and writes the
address block through it.
- sshd keeps its host keys under /data/forgefirm/ssh, so the fingerprint
survives updates; both sshd configs carry the same HostKey lines.
- forgefirm-logging passes logrotate a state file under /var/run
(logrotate refuses to run without one).
- forgefirm-persist points the boot timestamp and the random seed at
/data/forgefirm.
The dev image appends the /factory slot mounts, without nofail (busybox
mount hands it to the kernel, which rejects it). The rootfs command
entries lose their semicolons: on scarthgap the value is the task's
vardeps, split on whitespace, so "name;" left the function body out of
the signature and a changed body did not remake the rootfs; with the
bodies tracked, the dev image's DATETIME string needs a vardepsexclude.
release.sh gains the read-only gate (root ro, no /factory line,
ROOTFS_READ_ONLY=yes, host keys on /data). image.health checks the
mounts, the account binds, the banner bind, the host keys and the
dev-only /factory mounts.
Proven on the bench reference (dev image 20260909140901): / ro, /data
rw, /var/lib a tmpfs copy, the four account files and /etc/issue bound
from tmpfs, the host keys in /data/forgefirm/ssh, no "Read-only file
system" line in any log; forgectrl.auth and commission.account-login (a
temporary account rendered, logged in over HTTPS and removed again),
kernel.latch-locked-idle and motion.liveness-probe PASS; logrotate runs
with the volatile state. forgetest unit tests 335 OK; both images build
clean, and debugfs on the built rootfs shows every setting above.
forgectrl 6040e64 carries the three fixes behind the pinned 93fb22e: the
test header reached the way the sibling tests reach theirs, the key added
to /status carried in the panel dev-server mock, and the lens test's
carriage driven by the sweep's own steps rather than by a clock. The
shipping behavior is unchanged from 93fb22e; the pin names what passed.
Verified with bitbake -c fetch.
forgectrl 93fb22e fits the lens outcome text in the buffer the supervisor
shares with the probe. grblHAL-glowforge 97be92b takes the lens reference
from the realtime hook rather than settings-changed, so the controller no
longer overwrites the Z it just referenced.
Both verified with bitbake -c fetch at these revisions.
The Z session used to reference the lens itself with M103, which is gone:
the daemon sweeps the lens onto its hall edge before a controller starts
and leaves a marker, and the controller opens the Z envelope on that. No
daemon runs behind the harness, so nothing wrote the marker and every Z
move was refused, which is what the session's first move ran into.
The runner now writes the marker the daemon writes, and the session pins
lens_hall_edge_z_mm so the moves are counted from a known height: Z3 is 9
half-steps on the screw and Z4 is 12, so a 1 mm move up and back is 3
steps each way.
This is also the check that caught the controller overwriting its own
referenced Z at start, which the panel could not show.
forgectrl dd40dc8 takes the lens onto its hall edge in the motion-verify
window and gates the spawn when it cannot, and carries the per-axis
anchor, homed_axes on /status, and the panel reading Z from its own bit.
grblHAL-glowforge ef0f764 takes that reference as it loads its settings,
re-zeroes the kernel counters the daemon's GPIO steps never reached, and
drops M103.
Both verified with bitbake -c fetch at these revisions.
The lens now takes its hall-edge reference before any controller starts,
so Z is referenced on every start and M103 is gone. The laser-stream
harness opened its Z session by referencing the lens the way a
commissioning card did; it no longer has to, because Z is already open by
the time the session runs.
forgectrl.panel-serves gains the assertions for the per-axis reference:
homed_axes is an axis mask, homed agrees with it, and with a controller
running Z is referenced and reads inside the lens reach the same document
reports. That last check is the one that catches a panel showing nothing
for a Z the controller holds.
commission.check-motion already exercised the new path, because the
motion wizard's probe runs the same sequence the supervisor does, so its
covers map gains lenshome.c and its description names the lens reference
and the hard fault behind it.