cloud.verdict-hold printed to completion and then failed its hand-back on
head/z_mode=0. That job's header carried ZSmd 0, gfhardware's
set_mode_from_puls turns that into z_mode 0, and z_mode 0 is full step
(lenshome.c): the cloud client had set the lens the way the pulse file it
was playing asked. Nothing was left behind - the machine did what the job
said.
The baseline already skips nine attributes in cloud mode, for the reason
written above the list: the cloud client sets its own values for them
from every pulse header, and forcing the GRBL values under it would be
the baseline configuring another controller's machine. head/z_current and
head/z_mode are the same thing and were not on the list. They are now, in
cloud mode only; in GRBL mode the pair is still checked against 1 and 1,
which is the half of this an exemption could quietly break, so a test
holds it.
Host-proven: 89 baseline unit tests.
A leftover is dirt a run left behind. Four of the checks were reporting
something else: the machine part-way through work of its own, which it
finishes on its own a few seconds later. Each one failed a run that had
done nothing wrong.
laser.arm-wait-lid found it. The test passed every check it makes and
failed on cool=smoke/armed=False/hold=False. The smoke clear is the
engine's post-job work - run duty for smoke_s after an armed session,
timed, ending by itself (cool.c, Cool_Smoke) - and the machine was idle
fifteen seconds later without being asked. Armed and hold were both
false: nothing had been left.
The four, and what each was catching:
cool a phase alone (the smoke clear, a cooldown). Armed or
holding stays a leftover: the run left a job alive and the
engine is keeping the fans up for it.
controller the supervisor's own start. A takeover ends by starting
forgectrl again, and the respawn runs the liveness probe
and the lens reference before it reports running and
verified. The run had put it back.
state the ring draining to the end of a job. An underrun stays a
leftover and is still acknowledged with cnc/stop: that one
is the run's.
leds read_led read brightness, write_led writes target, and the
smooth trigger fades brightness toward target. An LED the
machine had already released still read lit mid-fade, and
the restore called itself done before the fade had moved.
Judged on target now: a run that left the button lit left
a target standing.
The waiting, the restoring and the reporting of anything a run actually
left are unchanged. What changed is which readings count as dirt.
Host-proven: 87 baseline unit tests, two of them new - a fading button
LED is not a leftover, a lit one is.
The airflow check reads the purge fan's off current with the fan off, and
the stand-down that follows commands it back on. The guard that proves
the machine was handed back whole read the draw immediately, so what it
got back was the off current the check had just measured: on the bench
reference, 74 against a 300 floor, with the fan drawing 631 a moment
later. The check failed for having worked.
The current follows the command; it does not arrive with it. The guard
now waits for the draw to reach the floor, up to PURGE_SPINUP_S, the way
the controller below it is already waited for, and logs how long it took.
It is no weaker: a fan that never reaches its floor still fails the
check, and the message now says how long it was given.
Pin: forgectrl 0.1.17 (a9e7073, the air assist and the purge in the
diagnostic run posture - the reason that check measured an idle fan and
wrote a floor from it).
The hand-back is the promise that a run leaves the machine where it found
it. Two of its checks reported instead of restoring, and the machine sat
in the state the run left it in.
The cooling engine: a run that ends without ending its job leaves one
alive behind it - the engine armed and holding for a job that is never
coming back, the fans at run duty. The check waited two minutes for that
to resolve itself, which it cannot, and wrote "failed: still
run/armed=True/hold=True". On the bench reference the fans then ran for
an hour.
The pulse ring: bytes the last job never played sit there, and the next
run replays them before its own. The check refused the return jog and
wrote "clear the ring (controller restart) before moving" - the
instruction, to a log, instead of the action.
Both end the same way: stop the controller and let the supervisor bring
it back. The job goes, the arm and the hold go with it, and the ring is
empty on the way in. stand_down() does that and proves it settled;
the cooling check calls it when the engine will not idle on its own, and
the return jog calls it when the ring has residue, refusing only if bytes
survive a restart. A leftover still reports what it found - the record of
what the run did is the point - but it reports having fixed it.
Host-proven: 62 baseline unit tests, including one that the hand-back
calls the stand-down for ring residue rather than describing it.
Run.snapshot() computed elapsed_s from the current time on every call,
whether the run had ended or not. The page shows the last run until the
next one starts, and it polls, so a finished run's figure went on
counting: the result badge said FAIL beside a number still climbing, and
the run read as still going. On the bench reference a test that ended
after 5082 s was showing 5242 s and rising.
The figure a finished run should carry was already recorded next to it:
finished["duration_s"], fixed when the result was written. snapshot()
now returns that once the run has ended, and the live count only while
it is running.
Host-proven: two unit tests - a finished run's clock reads its duration
and does not move across a poll, a running one's still climbs.
An empty form field never reaches the request, so POST /settings with an
empty value in the body carries no key at all and the daemon answers 400,
"no known setting in request". Only the query-string form gets an empty
value through. Measured on the bench reference:
POST /settings -d "lid_policy=" -> 400, the value unchanged
POST /settings?lid_policy= -> 200, the key cleared
Three places in the suite already knew this and say so in a comment; two
did not.
motion.lid-policy-hold restored with the form body, so on a machine whose
lid_policy has never been set the restore was refused and the test failed
its own check. It now uses the query-string form for the empty value, as
the others do. This is the same test as the commit before: that one
stopped it from writing the word "cancel" where the machine had nothing,
and this one lets it write the nothing.
cloud.mode-switch had the older shape of the same fault: it restored
homing_mode to the literal "none" where the machine had it unset. On the
bench reference the restore branch never ran, because that machine homes
through the cloud, so it has not fired yet; on a machine with homing_mode
unset it would have. It now restores exactly what it found.
The rest of the suite was audited for this and is clean: camera.key-read,
cooling.critical-tier, cooling.tec-drive, cooling.crash-watch-plumbing,
cooling.fire-watch-tiers, cooling.floor-and-warm-up, cloud.verdict-hold,
forgectrl.settings, logs.levels and motion.xy-microsteps all either guard
the empty value or go through the shared Restore helper, which has always
done it correctly. commission.cloud-header-capture uses that helper too.
motion.lid-policy-hold read the setting as
was = (fc.settings() or {}).get("lid_policy") or "cancel"
and wrote `was` back at the end. On a machine that has never set the
policy the setting reads as the empty string and behaves as cancel, so
the `or` turned "unset" into the word and the test handed the machine
back carrying a setting it did not arrive with. The hand-back reported
it, restored it, and failed the test.
The value is now captured exactly, empty included, and restored as
captured; the default belongs to reading the value, never to writing it
back. The read-back check gets the same default, so it no longer compares
None with the empty string. lid_policy_in_force carries the effective
policy into the evidence, which is what the old expression was reaching
for.
Found on the bench reference, on the run after the position dead band let
motion.lid-cancel-home through. The four other places in the suite that
read a setting with `or` are safe: two default to the empty string, which
is what unset is, and two feed a check rather than a restore. The shared
Restore helper captures raw values and writes the empty string back for
unset, as this now does.
Only this test's earlier passes are invalidated: the change is inside its
own body, and the per-test source hash of the other eleven motion tests
is unchanged (checked against the file before the edit).
The post-run pass compared the kernel position counters exactly, so a
test that put the head back within a hundredth of a millimeter failed
whenever that distance did not round to the same step.
motion.lid-cancel-home found it on the bench reference. The test cancels
a job with the lid twice and the controller returns the head to the job
start each time, landing 0.038 and 0.037 mm out - the same figures as the
run before it, which passed. This time the two returns left four steps on
X, 0.075 mm at 53.333 steps per mm, and the hand-back called it a
leftover. Twenty-three runs of that test, all passing, on returns of the
same accuracy: whether the residue rounds to zero is chance, not a
property of the machine or of the test.
A difference inside POSITION_DEADBAND_MM (0.1 mm, five steps at x8) on X
and Y is now the quantization rather than a leftover. The head is still
put back, so nothing accumulates over a campaign - only the failure goes.
Z stays exact: the return never moves the lens, so a Z difference is
still a leftover and still unrestorable.
Host-proven: four new unit tests on the boundary (four steps and a moved
Z on either side of it, and an unreadable reading), and the forgetest
suite.
The sheet handed the bench actuator every press after the operator's
presence press, and the actuator pressed every time - into nothing. On
the bench reference all six cards were pressed before the card asked:
the frame by 8 s, the focus by 33, the floor by 10, the dose curve by 8,
the corner by 9, and the flow-load card by 69. The operator then pressed
all six himself, which is the opposite of what the ready gate promises.
arm_press() waits for hw.button_lit(), which is true when any button LED
is on, and burn() started that wait before run_check had even started the
wizard. The button is lit through parts of a card that are not the arm -
the lens reference, the program on its way to the controller - so the
wait ended at once, the press landed before the job waited for it, and
the thread was gone by the time the real cue came. forgectrl uses the
same predicate but only inside the job's own sample callback, with the
tube still dark, where a lit button does mean the arm.
The machine already says when it wants the press: a live check opens a
`press` wait prompt at that moment, and run_check sees every prompt. It
now presses there, through a new Ctx.press_now() - no LED read, no
waiting thread, no timing guess. That retires the per-card lit-timeout
column of CARDS, which existed only to give the flow-load card's coolant
settle enough room for a wait that was reading the wrong thing.
The four grbl-driven arm_press() callers in laser.py and cooling.py are
left as they are: they call it with the machine idle and its LEDs dark,
so the level read is the edge they mean. The same shape would bite them
if that ever stopped being true.
Pin: forgectrl 0.1.15 (6e69479, the flow-load tail ending at the coolant
peak and the way out of a finished setup page).
Host-proven: 357 forgetest unit tests, including two new ones - the
actuator presses on the prompt with button_lit stubbed false throughout,
so the LED is provably not consulted, and the press falls to the operator
without a takeover.
One name for every machine was wrong: an operator with two of them on a
network had one forgefirm.local, and mDNS does not work on many networks
at all. The machine now calls itself forgefirm-<xxxx>, from the last four
hex digits of its WiFi MAC address, and sends that name with its DHCP
request, so a network with dynamic DNS publishes it and a router lists
the machine by name. The name is the same at every boot, two machines
take different names, and no serial number leaves the machine.
forgefirm-hostname (new): reads the wlan0 MAC address (eth0 on a machine
with no WiFi) at S38 in rcS, after udev has probed the network drivers
and before poky's hostname.sh reads the file and before the network
starts. The rootfs is read-only, so the name is written through a
bind-mounted copy under /run/forgefirm. A bounded wait covers a slow
probe. hostname:pn-base-files is "forgefirm": the name before S38, and
the fallback when no MAC address can be read.
avahi is deleted - the bbappend, the daemon configuration, the service
file, the image install and the distro block. The address is the way in
that works on every network, and the DHCP name covers the rest.
forgefirm-banner: the marker lines are gone. "# ForgeFIRM addresses" and
"# end" delimited the address block inside /etc/issue, and getty prints
every line of that file, so both markers were on the console. The script
now keeps the image's own text in a second copy under /run/forgefirm,
captured once per boot before the first write, and renders the whole
banner from it. The block is the addresses alone: no mDNS name.
forgefirm-image.bb: the ForgeFIRM mark, under the OpenGlow one the base
image carries, with the version on the mark's own last line,
right-justified to the mark's last column. The mark is written once and
rendered per reader, because /etc/issue is parsed by busybox getty (a
backslash or a percent sign starts an escape, so the art goes in with
every backslash doubled) while /etc/motd is written out as it is. Widths
are measured in columns, not bytes: the color sequences take no room on
the screen. /etc/issue.net stays unused - the machine tells a client that
has not logged in nothing.
Acceptance: commission.mdns-announce is replaced by
commission.machine-name, which checks the name against the MAC address,
the bind-mounted /etc/hostname, the DHCP client's hostname option, the
banner's addresses, and that no mDNS responder is on the image; it covers
nothing by design, like the test it replaces. forgectrl.auth gains the
own-name Host check and its refusal with a domain on it. image.health
checks the /etc/hostname mount and the version on the mark's last line in
both files. commission.ssh-until-reboot asserts there is no
pre-authentication banner. commission_dark's lens coverage widens to
src/lenshome.* so src/lenshome.h is covered; the lint is clean at 83
tests.
Pins: forgectrl 0.1.14 (9e5330f, the hostname certificate and the Host
rule), meta-openglow ced2af2 (the DHCP hostname option and the motd mark)
in the kas lock.
Proven on the bench reference, hot-deployed and rebooted (image
20260910000208 dev): hostname forgefirm-b00a from MAC 2c:6b:7d:0d:b0:0a,
live and in the bind-mounted file; the DHCP client running with
-x hostname:forgefirm-b00a; the console banner and the motd carrying both
marks with the version aligned to the mark's last column, no marker line
and no .local name; forgectrl regenerating its certificate for the new
name. Host tests: 357 forgetest unit tests, forgectrl clean under
-Werror, tls_test and sanitize_test.
The acceptance page shares theme.css with the panel byte for byte, and forgectrl's copy gained the setup header's Download logs button. CI caught the drift, which is what that check is for; the pinned revision decides, so the copy follows.
The baseline has always examined the machine after every run and recorded
what the run left behind. It did nothing else with it: the leftovers went
to the log and the evidence, and the test still reported PASS. So a check
could measure correctly, walk away with the machine in a state nobody
chose, and be recorded green.
That is how the purge fan came to be left off by the airflow check. The
leftover was not even watched, but had it been, it would have been noted
and the test would have passed anyway, and an operator would still have
met the airflow hold at their first fire.
A post-run leftover now fails the run. One the baseline put back fails it
too: the restore is the bench cleaning up after a defect, not the defect's
absence. The message names what was left.
The baseline watches the head as well as the motion side now: purge air
on, which is how the machine idles, and the lens motor at its hold current
in half step, which the lens checks and the sheet cards take and must hand
back. The airflow check proves the machine is whole rather than merely
measured: afterward the purge fan must read commanded-on and must draw
above the floor the check itself just wrote.
The pin takes forgectrl 0.1.13 (a2d73ef), which restores the idle posture
after a diagnostic, fixes the lens session's takeover flag, and gives the
setup a Download logs button, since the panel's Logs tab is unreachable
until the setup is complete.
Expect this to find things. A test that has been handing the machine back
imperfectly has been passing until now, and the first campaign under the
rule is where that shows.
The guard keeps scripts/manifest-from-tree.py and forgefirm-image-manifest.bbclass computing the same identity, so it pins the skip list by value and the new suffix failed it. It now expects the version file in both places, and a second case proves the point the change exists for: setting the release number leaves the layer hash where it was. My miss: I proved the identity by hand and pushed without running the unit suite first.
Setting the release number was a platform change. FORGEFIRM_RELEASE sat
in forgefirm-image.bb, the recipe hashes as content of meta-forgefirm,
and a change to the content of a layer invalidates every acceptance
result. So a version bump threw away the campaign that was meant to
authorize that very release, and the number therefore had to be decided
before the image the campaign ran on. Nothing said so: the release-flow
page went straight from the kas configuration to the artifact and the
pipeline, while the gate quietly required the recipe value, the rootfs
stamp, the archive's meta-version and the tag to agree. v0.0.1 was cut
on a tree whose number happened to be right; the next one would have
cost a second campaign to discover the rule.
The number moves to forgefirm-release.inc, which carries it and nothing
else, and the manifest leaves that file out of the layer content hash
exactly as it leaves out the component pin files
(FORGEFIRM_MANIFEST_VERSION_SUFFIX, and the same list in
scripts/manifest-from-tree.py, which computes the identity on a
workstation and must agree byte for byte). release.sh reads the number
from the new file.
The version is metadata, not platform content, and this only makes the
manifest say what it already meant: the version string was already
outside the identity hash, and it was the file carrying it that defeated
that. Nothing is weakened. release.sh still requires the number to equal
the rootfs stamp, the .fw meta-version and the release tag, and
image.health still compares the stamp on the running machine with the
manifest's.
Proven: the tree manifest is byte-identical across a bump from 0.0.1 to
0.0.2 (identity a64e51b8e5ecca0af683d4f0 either way, the meta-forgefirm
layer hash unchanged), where before the two differed. bitbake resolves
FORGEFIRM_RELEASE=0.0.1 and FORGEFIRM_VERSION_STRING=v0.0.1 for the
release image through the new require, and the dev image still overrides
the string with its build timestamp.
Campaign c-20260909204732-3fa9 on image 20260909193551, 83 tests, 83 satisfied (62 inherited), release authorized. It replaces the artifact of the release that was withdrawn: that one authorized an earlier rootfs, and the gate compares against the rootfs it is asked to sign.
ForgeFIRM's own files live under /data/forgefirm; two configuration
files did not. The machine settings sat at /data/forgefirm.conf, in the
root of /data beside the factory's own files, and the cloud-mode
configuration sat at /data/etc/gfhome.conf, inside a directory the
factory owns. Both move:
/data/forgefirm.conf -> /data/forgefirm/forgefirm.conf
/data/etc/gfhome.conf -> /data/forgefirm/gfhome.conf
There is no migration: only the bench has ever run this firmware.
/data/etc now holds only the factory's wpa_supplicant.conf.
The acceptance check of the file modes reads the settings file at its
new path, and the two bench tools that read it directly follow. The
pins move to the revisions that carry the change, forgectrl also
bringing the fix that reads the module's disabled state as idle:
forgectrl 468ee21 (0.1.12)
grblhal-glowforge 9ee624b (0.1.10)
forgefirm-app 56f134a (0.1.28+git)
The lock moves meta-openglow to b7ad6d9, which pins python3-gfhardware
on the same revision. The four upstream layers stay where they were:
`kas lock --update` moves every floating repository, and a release is
not the place to take poky, meta-openembedded and meta-freescale along
for the ride.
A machine out of the box sat in the module's power-on state, disabled, for the whole of its first run, and every idle gate read that as busy: the setup's sensors check refused to start, settings writes answered 409, and the cooling engine held cooldown airflow from boot. forgectrl now reads disabled as idle (the state means no program in progress), with fault and underrun still busy until acknowledged and an unreadable state still failing closed.
The yocto-cold-build workflow and its kas/ci.yml overlay built the release image on a hosted runner as a reproducibility probe. It never ran to completion, its first dispatch (2026-09-09) stopped on the runner's user-namespace rule, and a probe nobody runs is a trap. Every Yocto build, the release included, runs on the build host; the release proof is the local pipeline (release.sh) and the bench campaign. The pre-publish checklist loses its self-containment line to match.
BitBake isolates the network of its tasks with a user namespace, and the ubuntu-24.04 hosted runner's AppArmor profile refuses that to an unprivileged process, so the cold build stopped before its first task (run 34381825302). The workflow lifts the restriction for the run; nothing in the layers or the image changes.
The kas configuration takes the pinned-remote meta-openglow block, with
its commit in the lock file (d655e1e, the read-only rootfs), so a fresh
clone builds the release without a sibling checkout. The lock keeps the
upstream layers where they were.
releases/v0.0.1 carries the acceptance artifact the bench exported for
this image: campaign c-20260909160235-7649 on 20260909150456, 83 tests,
83 satisfied, none inherited, release authorized. The release gate
recomputes every test's fingerprint from the manifest inside the release
rootfs and signs only when the recorded results agree.
forgectrl now holds the motion check while a lid or the interlock is
open instead of starting the controller unverified: GET /mode reports
controller "waiting" with why, the button blinks amber, and the check
runs when the enclosure closes. motion.gate-waits-for-lid drives the
fixture's lid channel: the lid opens, forgectrl restarts, /mode must
read waiting with why naming the lid, no pid, motion unverified, the
button amber (sampled over a blink period: the smooth trigger's target
reads 0 through the off half) and no probe line in the log; the lid
closes, and the controller must come up verified with MOTION OK on the
first probe.
Proven on the bench reference: PASS, the controller verified 6.5 s
after the lid closed.
The rootfs mounted read-write, so a slot ran with its own files open to
change, and the factory-slot mounts rode along on the release image.
Both images now carry the read-only-rootfs feature: the ro root line and
the rcS default, the volatile links made at rootfs time, a writable copy
of /var/lib at boot, a build failure for a post-install that needs the
machine, and the removal of shadow, base-passwd, update-rc.d and
update-alternatives.
What must last or change at run time is handled file by file:
- forgefirm-users renders the four account files from the record into
/run/forgefirm/accounts and bind-mounts each copy over its /etc file
(useradd and the rest are gone with shadow); a render writes through
the mount, and the image's own files apply until the first render.
- forgefirm-banner bind-mounts a copy of /etc/issue and writes the
address block through it.
- sshd keeps its host keys under /data/forgefirm/ssh, so the fingerprint
survives updates; both sshd configs carry the same HostKey lines.
- forgefirm-logging passes logrotate a state file under /var/run
(logrotate refuses to run without one).
- forgefirm-persist points the boot timestamp and the random seed at
/data/forgefirm.
The dev image appends the /factory slot mounts, without nofail (busybox
mount hands it to the kernel, which rejects it). The rootfs command
entries lose their semicolons: on scarthgap the value is the task's
vardeps, split on whitespace, so "name;" left the function body out of
the signature and a changed body did not remake the rootfs; with the
bodies tracked, the dev image's DATETIME string needs a vardepsexclude.
release.sh gains the read-only gate (root ro, no /factory line,
ROOTFS_READ_ONLY=yes, host keys on /data). image.health checks the
mounts, the account binds, the banner bind, the host keys and the
dev-only /factory mounts.
Proven on the bench reference (dev image 20260909140901): / ro, /data
rw, /var/lib a tmpfs copy, the four account files and /etc/issue bound
from tmpfs, the host keys in /data/forgefirm/ssh, no "Read-only file
system" line in any log; forgectrl.auth and commission.account-login (a
temporary account rendered, logged in over HTTPS and removed again),
kernel.latch-locked-idle and motion.liveness-probe PASS; logrotate runs
with the volatile state. forgetest unit tests 335 OK; both images build
clean, and debugfs on the built rootfs shows every setting above.
forgectrl 6040e64 carries the three fixes behind the pinned 93fb22e: the
test header reached the way the sibling tests reach theirs, the key added
to /status carried in the panel dev-server mock, and the lens test's
carriage driven by the sweep's own steps rather than by a clock. The
shipping behavior is unchanged from 93fb22e; the pin names what passed.
Verified with bitbake -c fetch.
forgectrl 93fb22e fits the lens outcome text in the buffer the supervisor
shares with the probe. grblHAL-glowforge 97be92b takes the lens reference
from the realtime hook rather than settings-changed, so the controller no
longer overwrites the Z it just referenced.
Both verified with bitbake -c fetch at these revisions.
The Z session used to reference the lens itself with M103, which is gone:
the daemon sweeps the lens onto its hall edge before a controller starts
and leaves a marker, and the controller opens the Z envelope on that. No
daemon runs behind the harness, so nothing wrote the marker and every Z
move was refused, which is what the session's first move ran into.
The runner now writes the marker the daemon writes, and the session pins
lens_hall_edge_z_mm so the moves are counted from a known height: Z3 is 9
half-steps on the screw and Z4 is 12, so a 1 mm move up and back is 3
steps each way.
This is also the check that caught the controller overwriting its own
referenced Z at start, which the panel could not show.
forgectrl dd40dc8 takes the lens onto its hall edge in the motion-verify
window and gates the spawn when it cannot, and carries the per-axis
anchor, homed_axes on /status, and the panel reading Z from its own bit.
grblHAL-glowforge ef0f764 takes that reference as it loads its settings,
re-zeroes the kernel counters the daemon's GPIO steps never reached, and
drops M103.
Both verified with bitbake -c fetch at these revisions.
The lens now takes its hall-edge reference before any controller starts,
so Z is referenced on every start and M103 is gone. The laser-stream
harness opened its Z session by referencing the lens the way a
commissioning card did; it no longer has to, because Z is already open by
the time the session runs.
forgectrl.panel-serves gains the assertions for the per-axis reference:
homed_axes is an axis mask, homed agrees with it, and with a controller
running Z is referenced and reads inside the lens reach the same document
reports. That last check is the one that catches a panel showing nothing
for a Z the controller holds.
commission.check-motion already exercised the new path, because the
motion wizard's probe runs the same sequence the supervisor does, so its
covers map gains lenshome.c and its description names the lens reference
and the hard fault behind it.
The three components carry documentation-only commits: the README
becomes an index card, AGENTS.md lands, and the comments that pointed at
the retired BRINGUP.md now point at the documentation site.
The pins move because the tree is already past the image that
20260907224817 built: a comment in the kernel config fragment
glowforge.cfg changed the meta-glowforge-bsp layer content hash, which
the manifest counts as a platform change. Measured against that image's
manifest: meta-forgefirm and meta-openglow-core are unchanged,
meta-glowforge-bsp is not. A fresh image build and a full acceptance
campaign therefore precede any release, and holding these pins back
would buy nothing.
Verified: every pinned revision is on its public remote, and
bitbake -c fetch resolves all of them.
Every fact in the two documents is now on the documentation site, which
is the single source of truth. This repository carries no project
documentation any more: it is the build and release base plus the
acceptance tool, the bench tools and the fixture firmware.
BRINGUP.md was the runbook, the hardware facts bank and the open-work
list. CAMPAIGN-LOG.md was the dated record of how each result was
obtained. What replaces them: the site for present state, and the
commit message for the record of what a change did and how it was
proven, so the change and its record stay together. Local open work is
the developer's own file at the tree root and is not tracked here.
README.md becomes an index card: what this is, build, test, and where
the documentation is.
The release pipeline tags the documentation. Firmware on a machine
needs the documentation that agrees with it, so release.sh now tags the
forgefirm-docs checkout with the same v<version> as the release, and
prints the command that pushes the tag with the release. The checkout
must exist and be clean, which is a new gate before the signature.
FORGEFIRM_DOCS_DIR names the checkout (default: the sibling one) and
FORGEFIRM_DOCS_SKIP releases without a tag, loudly, and is never the
default. The tag is made at staging and pushed with the release, never
before: a documentation tag for a release that never shipped is worse
than no tag.
No catalog consequence. release.sh is host-side and is in no image.
The commission.py change is one sentence of a test description, not
behavior. accel_crash_probe.py and the kas header lose pointers to the
retired files.
Checks: bash -n and sh -n on release.sh, and the tracked trees carry no
reference to either retired file.
grblhal-glowforge e2ba043 carries the core fork at 362577d, whose block
ring index covers $398 up to 1000. Fetch-verified with bitbake -c fetch.
Pin bump only.
scripts/bench/planner_blocks_test.py restarts the null-sink controller
at $398=400 and at 1000 on one settings store and requires an answer on
the port, the depth in the status report and a move to Idle. It runs in
the grblHAL repo's CI; registered on the bench page and in the README.
BRINGUP: $398 runs over its whole range, the spin item is closed, and
the XY microstep item is committed, pushed and pinned. CAMPAIGN-LOG: the
landing and the index fix, with the host and bench proof.
No catalog consequence beyond the core submodule the motion tests
already cover; the coverage lint is clean.
forgectrl 0.1.6 (b1eee4d): the xy_microsteps setting and the quiet hold route.
grblhal-glowforge 0.1.5 (48d5f1d): the XY scale, tick and stop ramp derived
from the microstep mode, plus the producer lead ceiling.
forgefirm-app 0.1.25+git (d1c47b8): the checked header keys note; the
python3-gfhardware pin in meta-glowforge-bsp moves with it.
Fetch-verified with bitbake -c fetch. Pin bump only: the acceptance
consequence is the catalog invalidation the manifest already derives.
The baseline derives x/y_mode, step_freq, ramp_rate and the configured
markers from the xy_microsteps setting; ramp_rate joins the GRBL
controller's set (the driver writes it; the cloud client runs at the
module's). motion.microstep-modes cycles 8, 16 and 32: the save restarts
the idle controller, the kernel reads the mode with its tick and ramp,
$100/$101 are the mode's and a typed $100 is overwritten, a 40 mm jog at
top speed returns to Idle with the kernel counters over the mode agreeing
with the commanded travel and the accelerometer seeing the head move; the
setting is put back as found.
Bench tools: xy_mode_test.py (the null-sink harness the grblHAL CI runs),
raster_dry.py (a top-speed raster per mode), xy_pattern_accel.py (the
operator's pattern from home with the machine silent and the head
accelerometer listening), arc_tolerance_sweep.py (a $12 ladder on the 9
in circle: the chord rate the protocol loop feeds, about 300 a second, is
the ceiling, not the core), and the xymode and xycircle live drills.
BRINGUP carries the present state and the facts; CAMPAIGN-LOG the dated
record, including the planner-blocks spin (a $398 of 255 or more loops
forever at start, a core bug) and its recovery.
The open item "Idle before the kernel drains" asked to decide between
holding Idle until the kernel drains and stopping the continuation pads
growing the lag. Measured first: on the machine, four chained 50 mm jogs
against cnc/state give 171, 175, 177 and 176 ms after Idle. One queue
depth, flat. The growth the item described is gone, so neither driver
change was made and the item is closed.
Host side agrees and says where the mechanism lives. Stream bytes are the
time axis, so a dumped stream's length is how long the machine plays it:
chained jogs produce 35755 bytes each with no growth, and the churn
session holds at 64790 bytes at a producer lead of 2 or 10 ms. Only above
the lead ceiling does it inflate. Rule 17 now holds that stream to a
budget derived from the job rather than to a recorded number, so the
inflation regime cannot return unnoticed; it fails at 5899 and 8005 ms
with the ceiling lifted and passes at 2301 against 3800.
BRINGUP carried a second error. It said every forgectrl path that stops
the controller after motion waits for cnc/state idle. super.c says
outright that POST /controller/stop is not idle-gated, because it is also
the emergency lever, and safes the machine with cnc/stop and the latch
before the signal instead. The mode switch, the cooling gate and the
daemon shutdown do gate on machine_is_idle(). Both the tail figure and
the gating claim are corrected, and CAMPAIGN-LOG carries the measurements
and the decision not to hold Idle.
Acceptance: rule 17 is a host harness rule in the grblHAL repo's CI, not
a catalog case. The catalog is unchanged because no machine behavior is:
the driver change is a bound on an out-of-range knob.
A new scripts/bench python file must be entered in the bench registry as
well as in that directory's README: the registry is what the acceptance
page reads, and a file in neither is a tool nobody can find. The harness
joins the other two host-side CI harnesses, marked not a bench-page tool
for the same reason they are - it drives the host-built null-sink
controller, not the machine.
CAMPAIGN-LOG gets the dated record of the rebase onto core build
20260905: the settings-struct measurement taken before the machine was
touched, the harness and the two broken builds it was validated against,
and the bench results.
forgectrl carries the GRBL settings store into the data directory, which
the acceptance suite already asserts. grblhal-glowforge follows the core
off the deprecated settings-changed event and moves the core submodule to
the fork rebased onto build 20260905.
Both revisions fetch.
The Z soft limit belongs to the driver, not to $20. glowforge_homing.c
owns sys.work_envelope, sys.homed and sys.soft_limits for Z, because Z is
always referenced, to the lens hall edge or to where the lens stands, and
the core knows neither. The core recomputes both masks from the settings
and drops Z when it does: $20 clears the soft-limit mask inside its
setter, and a $13x write clears the homed bit for the axis as well.
z_envelope_test.py drives the null-sink controller over TCP and holds the
rule. It checks that an unreferenced Z is collapsed to where the lens
stands and refuses a move each way, that X and Y stay free so the
reassert is Z's alone, and that neither write frees Z. It restores $132.
The harness runs in the grblHAL repo's CI, next to the laser stream and
lifecycle harnesses.
The release build merges kas/source-bundle.yml, which turns on the Yocto
archiver: the upstream source of each recipe as upstream publishes it, the
patches with their series file, and the recipe with its includes. The
overlay adds tasks only, so the image manifest is unchanged and an
acceptance result still applies; proven on the build host, where the
archiver build and a plain rebuild of the same tree give the same
content_sha256.
scripts/source-bundle.py packs forgefirm-source-v<version>.tar.gz: the
archives, both license manifests, the license texts, the ForgeFIRM layers,
the kas configuration, the layer revisions and the build identity of the
image. What the bundle must hold comes from the image, not from a list in
the script: every recipe of license.manifest and image_license.manifest
whose license is in the include list must have an archive, or the release
stops with the recipe named. release.sh attaches the bundle and covers it
with sha256sums.txt; FORGEFIRM_SOURCE_SKIP=1 bypasses deliberately.
On the build host: 68 of the image's 111 recipes carry source, 313.7 MiB,
under the 2 GiB limit of a release asset.
No acceptance catalog consequence: the change is release tooling on the
build host and puts no file and no behavior on the machine. The host-side
proof is forgetest/tests/test_source_bundle.py, which holds the license
decision, the choice of archive and the refusal.
BRINGUP names the new path in the standalone start line and the
stored-settings note. forgectrl.panel-serves checks, in GRBL mode, that
the store is /data/forgefirm/EEPROM-glowforge.DAT and that nothing of it
remains at the top of /data. super.c is already in its covers map.
meta-forgefirm: the forgefirm-users init replays the account at boot;
sshd refuses root and empty passwords and runs only while the panel
turns it on; the release image keeps an empty root password for the
console; the console banner; avahi announces forgefirm.local; https in
libmicrohttpd and ulfius; the panel on 80 and 443; the license bundle on
the rootfs; release.sh checks the root policy on the built rootfs.
forgetest: the commission suites (commission, commission_dark,
commission_sheet: 23 cases); the runner turns cloud mode on with the
typed phrase for a test that declares it; the baseline's motor_lock is
0; the log-export test checks the bundle for the camera key; the record
helpers write bytes as given and join the daemon's paths as POSIX. The
stream harness gains rule 24: a hold verdict is held again after a
resume. Bench drills: lens_travel.py and lens_stop_accel.py.
Docs: BRINGUP carries the present state; CAMPAIGN-LOG carries the dated
record.
Its last finding, the request-body cap, ran on image 20260904131106 in
campaign c-20260904132654-d731. forgectrl.auth carried the case: a 4 MiB
body from a client with no token was refused with 403, and the daemon
answered /status in the same second. The campaign is 56 of 56,
authorized.
That was the only audit item left in Next work, so item 9 goes and the
list is items 1 to 8. The record of the campaign, and of the retirement,
is in CAMPAIGN-LOG. The audit file is deleted, as the 2026-07-03 and
2026-08-13 audits were before it.
The request-body cap is on the bench. Image 20260904131106 carried
forgectrl ff89288 through a build overlay, and campaign
c-20260904132654-d731 authorized it, 56 of 56.
The pin now says what the campaign ran.
Three items in Next work tracked the 2026-09-01 audit: the remediation
follow-through, the PIC readings, and the deferred six. All three are
done, on one image, and the campaign the release gate asks for passed on
it, so they belong in the record rather than the open list.
What remains of the audit is one finding, the request-body cap, which is
now fixed and host-proven and needs an image and a campaign. It takes
their place as item 9.
forgectrl.auth gains the case that guards it: an oversized body from an
unauthenticated client is refused and the daemon is still serving after.
forgectrl 8ef9509 at 0.1.2, grblhal-glowforge 2f5edee at 0.1.2, and the
forgefirm-app recipes at dd0ebf3, 0.1.23+git. These are the sources the
full campaign on image 20260903211413 passed against, now that their
repositories are pushed and the build no longer needs the local overlay.
The campaign the release gate asks for, run on one flash with nothing
inherited and nothing hot-deployed. The entry records how the pause test
passed, not only that it did: its emission trail drains to zero and stays
there, against the earlier run's second kernel run inside the pause, and
the difference between them is that the actuator made the presses.
Also recorded: the two harness faults found on the way, the actuator
dropping off the network twice and being asked for silently, and what
now goes in the record so a repeat is answerable.
BRINGUP moves to the campaign image and drops the campaign from the owed
list, leaving the push and the pin bumps.
The ready gate lives inside the arm-and-fire helper, and the kill drill
calls that helper twice, once for the expected stop and once for the
SIGKILL. So the presence gate asked the operator for a second press part
way through a test they had already proved themselves present for, with
the actuator standing by holding the presses. It is the only test in the
catalog with two ready gates.
The second gate now returns at once. Its setup line still goes up,
because the second half may want the scrap moved, but there is no press
to make.
The bench actuator is an ESP32 on wifi and it reports its own signal
strength, which reads -83 dBm here, close to where an association
starts dropping. It has vanished twice tonight and taken a live test
with it, and neither time did the record say anything a reader could
use: only that it was gone.
Both numbers now go into the run's evidence, so the next drop says
whether the link faded or the box restarted.
A live test asked a person to click Ready on a page and then make
timing-critical presses in the middle of a burning cut. That is how
tonight's pause test became unanswerable: the actuator had dropped off
the network, the harness fell back to the operator without saying so,
and afterwards nobody could tell a second press from the machine
resuming on its own.
Where an actuator is up and wired to the button, the ready gate now
takes a press on the machine's own button as the presence check, and the
actuator performs every press in that test. The button does nothing at
Idle, so the press is only a presence check, and the gate waits for the
release so it is never read as the arm press. With no actuator the
operator does the presses and answers on the page, as before.
An actuator lost after that takeover is now said out loud, in the log
and in the evidence, instead of quietly becoming a person's press.
The live-fire cue was four lines of machine-shaped prose. It is now what
a person needs: protection, exhaust, extinguisher, scrap, lid.
The whole catalog ran green, 56 of 56, but on a hot-deployed suite file,
so it authorizes nothing; the entry says so first. What it did was find
the one real defect, the pic-soc-load bound that decided on a count of
noise, and put the arm-acknowledgment fix under live fire: the emission
witness passed with its airflow witness finding nothing, on the same
instant button press that produced last night's burn.
Also recorded: the boot clock that never stepped, and the bench fixture
dropping off the network mid-queue with no address pinned to fall back
to.
The check decided on one count of noise. Measured over eleven runs on the
bench the settled reader moves 3 counts off the idle regime (once 2, once
4) against a control split of 6 (once 7), so a half-split bound sits
exactly on the median: two of those eleven runs failed while the machine
read identically to the nine that passed, and the queue stopped on one of
them.
The kernel's spin is its own load level, a couple of counts under a
Python spin, so a settled reader lands above the idle regime without
reaching the busy one. Asking it to reach halfway was asking for
something the mechanism does not promise. What the check has to catch is
a settle that overshoots and lets the conversion fall back to idle, which
reads as no move at all, and a third of the split catches that with a
count of margin either way. The split collapsing is still the primary
proof, unchanged above.
The record of what the first attended run really found and what it cost:
the verdict had no run-session identity, so the arm opened its window on
the verdict computed for the idle session before it. The build entry
lists the four components taken from local commits and the checks made
on the built images.
BRINGUP moves to the new campaign image and says why nothing inherits.
The cooling verdict now carries the engine's own armed flag, so the
stand-in engines in both null-sink harnesses publish it. The lifecycle
harness gains two cases: an engine that never takes the armed window
must produce a refused arm and no emission, and one that takes it a
couple of seconds late must produce a wait and then a normal arm. The
late case is the one that proves the controller keeps reading the
verdict while it is blocked in the arm; without that every job would
fail there.
The emission witness gains the bench form of the same rule: no sample
may show the laser firing while the cooling engine reports a phase that
runs the fans at their idle duty. That is what a burn with no airflow
looks like from the outside, and nothing in the catalog looked for it.
The record of the live-fire failure on 2026-09-03: a job fired for about a
second with idle airflow under a warm-up hold that arrived one tick later,
the leftover start-gate setting behind it, the engine's arm-time gate gap it
exposed, and what is owed before any further live fire.
The record of the campaign's unattended set on p32, the numbers the two new
kernel drills and the air-assist offset diagnostic gave under the settle,
and the forgetest restart that interrupted the last test, with the rule
that keeps it from recurring.
The kernel's spin is its own load level, a count or two under a Python
spin, so the third check no longer compares the settled level with the
Python-spin control; it requires the settled idle reader to have moved at
least half the control's split off the idle regime. Bench: PASS on image
20260903011655 (control split +6, settled split +1).
BRINGUP item 9: items 10 and 11 are on image 20260903011655; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p32 build and its checks.
The PIC16F1713 converts its inputs in a free-running loop (10 channels,
about 0.30 ms a loop) and a read returns the last conversion of that
channel; the count follows the SoC's load at conversion time (idle 659,
busy 665 on the coolant thermistors, both tight; every channel shifts in
proportion to its count; the step lands within one PIC loop of the CPU
changing state, with a regulator's overshoot in each direction). The
kernel.pic-soc-load drill replaces kernel.pic-pacing: 200 reads after 3 ms
of sleep and 200 after 3 ms of spinning, with the module's settle off
(the control, reported) and on (the claim: the two agree). The catalog
counts 56 tests, 0 uncovered.
BRINGUP item 10 and the facts bullet describe the mechanism and the fix;
the CAMPAIGN-LOG entries record the first pass of the campaign, the PIC
study, and the mechanism's proof.
The sysfs state of watchdog0 says whether a process holds the device, and
none does: the kernel's core feeds the boot-armed hardware. The check
reads WCR through /dev/mem (WDE set, a 60 s period) and expects the state
to read inactive. Bench: PASS on image 20260903003213 with WCR 0x771f. The
BRINGUP facts bullet says the same.
BRINGUP item 9: items 10 and 11 are on image 20260903003213; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p31 build and its checks.
kernel.pic-pacing reads a coolant thermistor twice back to back, 300
pairs, with the module's pacing off (the control, reported) and on (the
claim: the second read agrees with the first). The bench proof for the
module's pic_gap_us pacing; the catalog counts 56 tests, 0 uncovered.
BRINGUP item 10 and the facts bullet describe the pacing as it is; the
CAMPAIGN-LOG entry records the change and its proof.
BRINGUP item 9: the deferred batch is on image 20260903000529; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p30 build and its checks.
B-16: BB_SIGNATURE_LOCAL_DIRS_EXCLUDE in the distro conf names __pycache__
and .pytest_cache, so a workstation's bytecode caches never enter a
file:// checksum (proven in the build VM: a cache under the package leaves
the fetch task alone, a source change reruns it).
kernel.resume-lead: two phases behind one takeover at a 1 kHz tick. E, a
resume whose lead is longer than the data ends at end-of-data within a few
ticks (a lost end-of-data would show as 255 ms). L, a 1000-byte lead over
FIRE bits with the latch unlocked and the chain unarmed keeps the FIRE
line low through the lead and drives it from the waypoint byte on. The
bench proof for the module's K-4 and K-8 fixes; the catalog counts 55
tests, 0 uncovered.
BRINGUP items 9 and 11 and the CAMPAIGN-LOG entry record the batch and the
campaign rule: no campaign until every audit finding is on one image.
CAMPAIGN-LOG: the push in CI order, the pin bumps, and build p29 with its
checks. BRINGUP item 9: the remediation is pushed, pinned, and built; owed
now is the flash, the fresh-boot baseline, and the full campaign on the
image, which the release gate asks for.
B-16: the forgetest recipe fetches the package directory whole, and the
file fetcher's checksum is taken before do_unpack drops __pycache__, so a
host `__pycache__` written by running the tests moves the recipe's task
hash with no source change. The three python steps run with -B now;
`src-sync` to the build VM excludes the caches as well. A developer who
runs the tests locally without -B still creates them; there is no exclude
on the file fetcher, so this is the floor, not a cure (BRINGUP item 11).
The audit remediation's commits and the bench session's fixes in the
three repositories, pushed and bench-proven on image 20260902144848
(built from the same commits through a local pin overlay). Each PV
moves with its SRCREV; the manifest records the PV.
The hardware watchdog is armed by the bootloader and kept fed by the
kernel core; nothing in userspace opens it. Until now nothing in the
catalog read that it is running. image.health reads watchdog0's state,
timeout and bootstatus (CONFIG_WATCHDOG_SYSFS, meta-openglow) and
requires active at 60 s; bootstatus goes into the evidence, so a campaign
that follows a watchdog reset says so.
Bench: on 2026-09-02 the forced hang ended in the factory recovery after
the 60 s timeout; the register read WCR 0x771f. The sysfs view itself
arrives with the next image, so this check is proven there.
Catalog consequence: image.health is always-run.
CAMPAIGN-LOG: the forced hang on image 20260902144848 with its times,
the watchdog register read back, U-Boot's recovery branch (the purple
button), the recovery image on the lease, and the power cycle back; the
pad measurement dropped by decision. BRINGUP item 9: both measurements
closed; the pushes, the pins, one build and the campaign on that image
remain.
CAMPAIGN-LOG: the attended set green on image 20260902144848, the three
defects the live runs and the takeover restarts found (the dwell-gap
latch rule, the log drops, the busy-start airflow) and their fixes, and
what is owed. BRINGUP item 9: both queues green; the measurements, the
pushes, the pins, one build, and the campaign on that image are what is
left.
A forgectrl started while the kernel is not idle takes the cooldown
airflow (forgectrl's busy-start rule), and on the bench it kept it: after
kernel.fire-line's takeover restarted the daemon with the kernel in the
drill's safe state, the exhaust ran at 6200 rpm on an idle machine until
the daemon was restarted by hand. cooling.fans-quiet-after-motion gains
the case: forgectrl stopped, cnc/disable written, forgectrl started, and
within 90 s the controller must be running with the idle duties applied.
The host replay stubs the init script and holds both outcomes: the idle
duties after the start, and a daemon that keeps the cooldown duties.
Bench: with forgectrl 522cdb2, the busy start logged, idle airflow one
tick later, the duties idle 15 s after the start, PASS.
Catalog consequence: the cooling.* implementation hashes move.
laser.emission-witness required the hardware button latch clear in every
sample the engine reported armed, and, after a first fix, in every sample
up to the last nonzero emission count. Both windows were drawn from
lagging signals: the engine's armed flag follows the controller's next
report, and the emission counter latches once per second and reads
nonzero about two seconds past the relock. Both reached into the tail
where the job-end relock sets the button latch by design, and the rule
refused three clean runs on image 20260902144848 (all four sides
burned; the trail shows the latch clear from the press to the relock,
emission through the fourth side, HV_ENABLE's dip in the dwell and its
return).
The rule now uses the window the hardware defines: from the first
emission, in every sample whose readback word shows the laser latch
unlocked, the button-latch bit of that same word must be clear. That
spans the kernel-run gap of the dwell and ends at the relock, and no
lagging flag can misplace it. dwell_gap() is a pure function;
tests/test_laser_dwell.py holds the relocked tail, a set inside the gap,
and a trail without emission. The recorded trail of the third run
replays to a pass (47 unlocked samples, none set).
The live runs keep a per-sample trail in the evidence (TRAIL_FIELDS: the
readback word, the switches, the lock flag, the controller's state and
messages), so a run's timeline can be read back without a rerun.
A fourth run then errored on a name the refactor had removed and one
later check still used; py_compile does not catch it and a live drill
never executes on the host, so the CI job now fails on any undefined
name in the harness (pyflakes).
Catalog consequence: the laser implementation hashes move.
The daemons log to /dev/log with non-blocking datagrams and drop what
the socket will not take. The kernel's default queue for a unix datagram
socket is 10; forgectrl's arm-time burst alone was about 20, so the
lines around every job start were lost, and the daemon reported 18 to
46 dropped per job on the bench. The logging init sets
net.unix.max_dgram_qlen to 512 before rsyslog and the daemons start;
nothing on the image applied sysctl files before.
Bench: set at runtime on 20260902144848 for the rest of the session.
Catalog consequence: a layer change; the meta-forgefirm content hash
moves and nothing inherits on the image that ships it.
CAMPAIGN-LOG: the unattended set on image 20260902144848, the six
harness and diagnostic defects it found and their fixes, the numbers,
and the board's state at the end of the pass. BRINGUP: item 9 names the
reset in the hang case and the session under way; the facts bank gains
the PIC read-pattern measurement; Next work item 10 is a pacing of PIC
reads in the kernel.
update.slots-and-signature's apply section required 200 from
POST /update/apply, where the daemon answers 202 with started, like
every job endpoint, so its first bench run on image 20260902144848
ended before the job did; the cleanup then deleted the staged archive
under the running job. The drill requires 202 and started, and looks
for the daemon's refusal ("archive is not signed with the ForgeFIRM
release key"). Bench: the apply started, the job ended with that
refusal, PASS.
A queue started 4 s after a forgetest restart ran 7 tests instead of
10. The bench page's /state poll had a fixture probe in flight (an mDNS
answer), probe_fixture stamped its time at its start, and the queue
start read the stale fixture, none, so the three operator tests the
fixture runs in the unattended queue were routed to nobody. The probe
now runs under a lock and is stamped when it completes: a caller that
arrives during a probe waits for its answer. tests/test_fixture.py
holds the race with a slow scripted probe; it fails on the old code.
Catalog consequence: the update implementation hash moves; the runner
change is dev-only.
A queue runs the catalog in registration order among tests with the
same prerequisites, and cooling.aa-offset-calibrate followed
cooling.flow-verify. A flow-verify trial heats the tube water (a no-flow
trial by 17 C on the bench) and the warm slug circulates past the
coolant sensors for minutes afterward; the calibration's stationary gate
passed 44 s after the trial on image 20260902144848 and the edges read
the wave as disagreement.
The calibration is registered first now, with the reason beside it, and
tests/test_cooling_order.py holds the order.
Catalog consequence: the cooling.* implementation hashes move (the suite
file changed).
The hang case resumed the controller and sent $X alone. The stream
fault raises Alarm 17 (motor fault), which the core treats as a critical
event: $X is refused with error:79 until a soft reset, and the reset is
what the stream takes as the operator's acknowledgment of the fault
(grblHAL 38b450e: the kernel stopped and re-armed, the stale ring
cleared, the producer armed again). The drill's first bench run on image
20260902144848 therefore ended in Alarm.
The drill now records $X before the reset and requires the error:79
refusal, sends the soft reset, unlocks, requires Idle, requires the ring
back at its idle free count (the stale bytes of the interrupted move are
gone), and then jogs. The final assertion read the state off the report
dict returned by wait_idle as a string; it never ran before because the
TIMEOUT test short-circuited it.
Bench: motion.deadman PASS on 20260902144848 (kill respawn 1.2 s, hang
to underrun 0.21 s, $X -> ALARM:17 error:79, reset + $X -> Idle, ring
33521664 of 33521664, jog Jog -> Idle, restart retook supervision).
Catalog consequence: the motion.* implementation hashes move (the suite
file changed) and the set re-ran and passed. The failed first run had
closed the campaign, so the always-required core ran again in the new
one, as the campaign rules require.
The gate that refuses a latch unlock while the chain may hold HV_ENABLE
up (charge_pump_alive or a pulse engine not idle) ran at the start of
phases B, U and K3 of kernel.fire-line, within a second of the previous
phase's run. A run feeds the charge-pump watchdog every 200 ms and the
one-shot holds ALIVE for 0.45 s after the last feed, so the gate read
alive=1 and refused: the first bench run of the gate (forgefirm
64f552fc; the laser_pgood gate before it was vacuous) failed phase B on
image 20260902144848.
wait_hv_off() polls the chain for up to 3 s before it refuses, logs the
release when it was not immediate and records every wait in the
evidence (hv_release_s). require_hv_off and check_hv_off use it. The
bench scripts that copy the gate (fire_test.py per phase,
gate_a_kernel_drills.py K3 after K2) get the same wait.
Bench: kernel.fire-line PASS on 20260902144848 with the chain released
after 0.41 s at each of the three phase boundaries. Host:
tests/test_kernel_suite.py covers release inside the window, a chain
held past it, and a chain already off.
Catalog consequence: the kernel.* implementation hashes move (the suite
file changed); the kernel set re-ran and passed.
The manifest lists the modules directory without the kernel's
LOCALVERSION_AUTO hash (the hash does not reproduce across a re-patch of
the same source, so the image manifest strips it). image.health still
compared the full running release against that list and failed on the
first post-flash run of image 20260902144848 with the kernel
6.12.20-fslc-fslc-g72a0b1431a9d against the manifest's 6.12.20-fslc-fslc.
kernel_ident() strips the same suffix from both sides, so a manifest with
or without the hash matches the running kernel, and a different base
release still fails. tests/test_image.py covers both forms.
Catalog consequence: only image.health's own implementation hash moves;
it is an always-run test, so no inherited result is affected.
BRINGUP describes the present: the 54-test catalog and its seven-test
always core, the tier counts, the shipped low-temperature gates, the
density floor ($35 = 10), the two local core commits, the ffboot env
write, the aa-offset route, the current bench image, and the bench
measurements the audit asks for (pooled into the next session). The
workstation shell notes and every em dash are gone.
forgetest: the takeover waits for the cloud client too (found by its
command line); the unauthenticated /boot probe names the endpoint's
parameter; the UI prose is American English. Recipes: forgetest
fetches its package directory and init script only and drops
__pycache__ at unpack; the dev image no longer re-adds forgectrl; the
release image's remove list drops the gfui-client the BSP no longer
has; the platform identity strips the kernel's local-version hash
from the modules directory name, so a re-patched kernel keeps its
fingerprints. grblhal restart is stop then start. release.sh --dev
packs the dev image. fixture.sh refuses a readable env file.
Bench tools: the live-fire drills measure the lid-IR baseline before
every run and point at the fire-watch thresholds the engine reads;
one thermistor conversion (gfbench.degc) serves every drill; the six
dated measurement records leave the tool directory; feeder.c names the
two sysfs writes its caller makes.
Host tests: forgetest 258 pass; the coverage lint reports no uncovered
path across 54 tests. Acceptance: forgectrl.auth covers the /boot
probe; update.* cover ffboot and the manifest identity; the runbook
and bench-tool changes have no catalog consequence.
A read-only mount of an ext4 slot still replays its journal, which
writes to the partition the probe was not meant to touch. The probe
mounts with noload: the journal is left as it is and the slot's
content is read as it stands.
Acceptance: update.slots covers ffboot -l (covers map ffboot/**).
A test's fingerprint covered its own text and its module's shared text
only, so a judge imported from a sibling suite module (laser.py takes
its motion judges from motion.py) could change without moving the
fingerprints of the tests that call it. The shared text of every sibling
module a module imports now rides along, transitively; unit test.
update.slots-and-signature claimed to refuse a tampered signature but
fed fwup one garbage file. It now makes a throwaway key pair on the
machine, signs a tiny archive, checks that the archive verifies with its
own key and fails against the shipped release key, and asks the update
job to apply it without confirm_unsigned: the job refuses it for its
signature before touching the slot.
BRINGUP named nine cool_* tunables (there are thirty), a cool_fire_ir_delta
key nothing reads, and panel source files that no longer exist; the
inventory is the present one. The bench page's help popovers pointed at a
documentation host and paths that do not exist; they open the site.
FORGEFIRM_RELEASE said 0.1.0, the first non-beta number by the settled
rule, so the first cut was either refused as 0.0.1 or shipped as a
non-beta. The recipe now says 0.0.1, and release.sh refuses a version at
or above 0.1.0 while the README carries the beta banner.