Commit Graph
100 Commits
Author SHA1 Message Date
ScottW514 b94c990bbc Leave the lens to the cloud client in cloud mode
cloud.verdict-hold printed to completion and then failed its hand-back on
head/z_mode=0. That job's header carried ZSmd 0, gfhardware's
set_mode_from_puls turns that into z_mode 0, and z_mode 0 is full step
(lenshome.c): the cloud client had set the lens the way the pulse file it
was playing asked. Nothing was left behind - the machine did what the job
said.

The baseline already skips nine attributes in cloud mode, for the reason
written above the list: the cloud client sets its own values for them
from every pulse header, and forcing the GRBL values under it would be
the baseline configuring another controller's machine. head/z_current and
head/z_mode are the same thing and were not on the list. They are now, in
cloud mode only; in GRBL mode the pair is still checked against 1 and 1,
which is the half of this an exemption could quietly break, so a test
holds it.

Host-proven: 89 baseline unit tests.
2026-09-10 13:08:33 -04:00
ScottW514 8d8ca2a8c6 Judge a hand-back on what a run left, not on the machine still working
A leftover is dirt a run left behind. Four of the checks were reporting
something else: the machine part-way through work of its own, which it
finishes on its own a few seconds later. Each one failed a run that had
done nothing wrong.

laser.arm-wait-lid found it. The test passed every check it makes and
failed on cool=smoke/armed=False/hold=False. The smoke clear is the
engine's post-job work - run duty for smoke_s after an armed session,
timed, ending by itself (cool.c, Cool_Smoke) - and the machine was idle
fifteen seconds later without being asked. Armed and hold were both
false: nothing had been left.

The four, and what each was catching:

  cool         a phase alone (the smoke clear, a cooldown). Armed or
               holding stays a leftover: the run left a job alive and the
               engine is keeping the fans up for it.
  controller   the supervisor's own start. A takeover ends by starting
               forgectrl again, and the respawn runs the liveness probe
               and the lens reference before it reports running and
               verified. The run had put it back.
  state        the ring draining to the end of a job. An underrun stays a
               leftover and is still acknowledged with cnc/stop: that one
               is the run's.
  leds         read_led read brightness, write_led writes target, and the
               smooth trigger fades brightness toward target. An LED the
               machine had already released still read lit mid-fade, and
               the restore called itself done before the fade had moved.
               Judged on target now: a run that left the button lit left
               a target standing.

The waiting, the restoring and the reporting of anything a run actually
left are unchanged. What changed is which readings count as dirt.

Host-proven: 87 baseline unit tests, two of them new - a fading button
LED is not a leftover, a lit one is.
2026-09-10 12:51:05 -04:00
ScottW514 4fbf14195b Wait for the purge fan's draw, do not read it in the same second
The airflow check reads the purge fan's off current with the fan off, and
the stand-down that follows commands it back on. The guard that proves
the machine was handed back whole read the draw immediately, so what it
got back was the off current the check had just measured: on the bench
reference, 74 against a 300 floor, with the fan drawing 631 a moment
later. The check failed for having worked.

The current follows the command; it does not arrive with it. The guard
now waits for the draw to reach the floor, up to PURGE_SPINUP_S, the way
the controller below it is already waited for, and logs how long it took.
It is no weaker: a fan that never reaches its floor still fails the
check, and the message now says how long it was given.

Pin: forgectrl 0.1.17 (a9e7073, the air assist and the purge in the
diagnostic run posture - the reason that check measured an idle fan and
wrote a floor from it).
2026-09-10 11:59:47 -04:00
ScottW514 d3fe1d90b9 Hand the machine back, do not describe what is wrong with it
The hand-back is the promise that a run leaves the machine where it found
it. Two of its checks reported instead of restoring, and the machine sat
in the state the run left it in.

The cooling engine: a run that ends without ending its job leaves one
alive behind it - the engine armed and holding for a job that is never
coming back, the fans at run duty. The check waited two minutes for that
to resolve itself, which it cannot, and wrote "failed: still
run/armed=True/hold=True". On the bench reference the fans then ran for
an hour.

The pulse ring: bytes the last job never played sit there, and the next
run replays them before its own. The check refused the return jog and
wrote "clear the ring (controller restart) before moving" - the
instruction, to a log, instead of the action.

Both end the same way: stop the controller and let the supervisor bring
it back. The job goes, the arm and the hold go with it, and the ring is
empty on the way in. stand_down() does that and proves it settled;
the cooling check calls it when the engine will not idle on its own, and
the return jog calls it when the ring has residue, refusing only if bytes
survive a restart. A leftover still reports what it found - the record of
what the run did is the point - but it reports having fixed it.

Host-proven: 62 baseline unit tests, including one that the hand-back
calls the stand-down for ring residue rather than describing it.
2026-09-10 11:34:27 -04:00
ScottW514 594b6990fd Stop a finished run's clock
Run.snapshot() computed elapsed_s from the current time on every call,
whether the run had ended or not. The page shows the last run until the
next one starts, and it polls, so a finished run's figure went on
counting: the result badge said FAIL beside a number still climbing, and
the run read as still going. On the bench reference a test that ended
after 5082 s was showing 5242 s and rising.

The figure a finished run should carry was already recorded next to it:
finished["duration_s"], fixed when the result was written. snapshot()
now returns that once the run has ended, and the live count only while
it is running.

Host-proven: two unit tests - a finished run's clock reads its duration
and does not move across a poll, a running one's still climbs.
2026-09-10 10:43:46 -04:00
ScottW514 57ba3454b4 Clear a setting the only way the daemon accepts, everywhere it is cleared
An empty form field never reaches the request, so POST /settings with an
empty value in the body carries no key at all and the daemon answers 400,
"no known setting in request". Only the query-string form gets an empty
value through. Measured on the bench reference:

    POST /settings -d "lid_policy="   -> 400, the value unchanged
    POST /settings?lid_policy=        -> 200, the key cleared

Three places in the suite already knew this and say so in a comment; two
did not.

motion.lid-policy-hold restored with the form body, so on a machine whose
lid_policy has never been set the restore was refused and the test failed
its own check. It now uses the query-string form for the empty value, as
the others do. This is the same test as the commit before: that one
stopped it from writing the word "cancel" where the machine had nothing,
and this one lets it write the nothing.

cloud.mode-switch had the older shape of the same fault: it restored
homing_mode to the literal "none" where the machine had it unset. On the
bench reference the restore branch never ran, because that machine homes
through the cloud, so it has not fired yet; on a machine with homing_mode
unset it would have. It now restores exactly what it found.

The rest of the suite was audited for this and is clean: camera.key-read,
cooling.critical-tier, cooling.tec-drive, cooling.crash-watch-plumbing,
cooling.fire-watch-tiers, cooling.floor-and-warm-up, cloud.verdict-hold,
forgectrl.settings, logs.levels and motion.xy-microsteps all either guard
the empty value or go through the shared Restore helper, which has always
done it correctly. commission.cloud-header-capture uses that helper too.
2026-09-10 10:26:25 -04:00
ScottW514 d58ee39086 Hand lid_policy back unset where the machine had it unset
motion.lid-policy-hold read the setting as

    was = (fc.settings() or {}).get("lid_policy") or "cancel"

and wrote `was` back at the end. On a machine that has never set the
policy the setting reads as the empty string and behaves as cancel, so
the `or` turned "unset" into the word and the test handed the machine
back carrying a setting it did not arrive with. The hand-back reported
it, restored it, and failed the test.

The value is now captured exactly, empty included, and restored as
captured; the default belongs to reading the value, never to writing it
back. The read-back check gets the same default, so it no longer compares
None with the empty string. lid_policy_in_force carries the effective
policy into the evidence, which is what the old expression was reaching
for.

Found on the bench reference, on the run after the position dead band let
motion.lid-cancel-home through. The four other places in the suite that
read a setting with `or` are safe: two default to the empty string, which
is what unset is, and two feed a check rather than a restore. The shared
Restore helper captures raw values and writes the empty string back for
unset, as this now does.

Only this test's earlier passes are invalidated: the change is inside its
own body, and the per-test source hash of the other eleven motion tests
is unchanged (checked against the file before the edit).
2026-09-10 10:13:16 -04:00
ScottW514 b2c42f0cc3 Do not fail a hand-back on the step the counters round to
The post-run pass compared the kernel position counters exactly, so a
test that put the head back within a hundredth of a millimeter failed
whenever that distance did not round to the same step.

motion.lid-cancel-home found it on the bench reference. The test cancels
a job with the lid twice and the controller returns the head to the job
start each time, landing 0.038 and 0.037 mm out - the same figures as the
run before it, which passed. This time the two returns left four steps on
X, 0.075 mm at 53.333 steps per mm, and the hand-back called it a
leftover. Twenty-three runs of that test, all passing, on returns of the
same accuracy: whether the residue rounds to zero is chance, not a
property of the machine or of the test.

A difference inside POSITION_DEADBAND_MM (0.1 mm, five steps at x8) on X
and Y is now the quantization rather than a leftover. The head is still
put back, so nothing accumulates over a campaign - only the failure goes.
Z stays exact: the return never moves the lens, so a Z difference is
still a leftover and still unrestorable.

Host-proven: four new unit tests on the boundary (four steps and a moved
Z on either side of it, and an unreadable reading), and the forgetest
suite.
2026-09-10 09:53:27 -04:00
ScottW514 1dd6ad608b Press the button when the machine asks, not when its LED is on
The sheet handed the bench actuator every press after the operator's
presence press, and the actuator pressed every time - into nothing. On
the bench reference all six cards were pressed before the card asked:
the frame by 8 s, the focus by 33, the floor by 10, the dose curve by 8,
the corner by 9, and the flow-load card by 69. The operator then pressed
all six himself, which is the opposite of what the ready gate promises.

arm_press() waits for hw.button_lit(), which is true when any button LED
is on, and burn() started that wait before run_check had even started the
wizard. The button is lit through parts of a card that are not the arm -
the lens reference, the program on its way to the controller - so the
wait ended at once, the press landed before the job waited for it, and
the thread was gone by the time the real cue came. forgectrl uses the
same predicate but only inside the job's own sample callback, with the
tube still dark, where a lit button does mean the arm.

The machine already says when it wants the press: a live check opens a
`press` wait prompt at that moment, and run_check sees every prompt. It
now presses there, through a new Ctx.press_now() - no LED read, no
waiting thread, no timing guess. That retires the per-card lit-timeout
column of CARDS, which existed only to give the flow-load card's coolant
settle enough room for a wait that was reading the wrong thing.

The four grbl-driven arm_press() callers in laser.py and cooling.py are
left as they are: they call it with the machine idle and its LEDs dark,
so the level read is the edge they mean. The same shape would bite them
if that ever stopped being true.

Pin: forgectrl 0.1.15 (6e69479, the flow-load tail ending at the coolant
peak and the way out of a finished setup page).

Host-proven: 357 forgetest unit tests, including two new ones - the
actuator presses on the prompt with button_lit stubbed false throughout,
so the LED is provably not consulted, and the press falls to the operator
without a takeover.
2026-09-10 08:41:53 -04:00
ScottW514 f0c40e7d4f Name every machine after its own MAC address, and drop mDNS
One name for every machine was wrong: an operator with two of them on a
network had one forgefirm.local, and mDNS does not work on many networks
at all. The machine now calls itself forgefirm-<xxxx>, from the last four
hex digits of its WiFi MAC address, and sends that name with its DHCP
request, so a network with dynamic DNS publishes it and a router lists
the machine by name. The name is the same at every boot, two machines
take different names, and no serial number leaves the machine.

forgefirm-hostname (new): reads the wlan0 MAC address (eth0 on a machine
with no WiFi) at S38 in rcS, after udev has probed the network drivers
and before poky's hostname.sh reads the file and before the network
starts. The rootfs is read-only, so the name is written through a
bind-mounted copy under /run/forgefirm. A bounded wait covers a slow
probe. hostname:pn-base-files is "forgefirm": the name before S38, and
the fallback when no MAC address can be read.

avahi is deleted - the bbappend, the daemon configuration, the service
file, the image install and the distro block. The address is the way in
that works on every network, and the DHCP name covers the rest.

forgefirm-banner: the marker lines are gone. "# ForgeFIRM addresses" and
"# end" delimited the address block inside /etc/issue, and getty prints
every line of that file, so both markers were on the console. The script
now keeps the image's own text in a second copy under /run/forgefirm,
captured once per boot before the first write, and renders the whole
banner from it. The block is the addresses alone: no mDNS name.

forgefirm-image.bb: the ForgeFIRM mark, under the OpenGlow one the base
image carries, with the version on the mark's own last line,
right-justified to the mark's last column. The mark is written once and
rendered per reader, because /etc/issue is parsed by busybox getty (a
backslash or a percent sign starts an escape, so the art goes in with
every backslash doubled) while /etc/motd is written out as it is. Widths
are measured in columns, not bytes: the color sequences take no room on
the screen. /etc/issue.net stays unused - the machine tells a client that
has not logged in nothing.

Acceptance: commission.mdns-announce is replaced by
commission.machine-name, which checks the name against the MAC address,
the bind-mounted /etc/hostname, the DHCP client's hostname option, the
banner's addresses, and that no mDNS responder is on the image; it covers
nothing by design, like the test it replaces. forgectrl.auth gains the
own-name Host check and its refusal with a domain on it. image.health
checks the /etc/hostname mount and the version on the mark's last line in
both files. commission.ssh-until-reboot asserts there is no
pre-authentication banner. commission_dark's lens coverage widens to
src/lenshome.* so src/lenshome.h is covered; the lint is clean at 83
tests.

Pins: forgectrl 0.1.14 (9e5330f, the hostname certificate and the Host
rule), meta-openglow ced2af2 (the DHCP hostname option and the motd mark)
in the kas lock.

Proven on the bench reference, hot-deployed and rebooted (image
20260910000208 dev): hostname forgefirm-b00a from MAC 2c:6b:7d:0d:b0:0a,
live and in the bind-mounted file; the DHCP client running with
-x hostname:forgefirm-b00a; the console banner and the motd carrying both
marks with the version aligned to the mark's last column, no marker line
and no .local name; forgectrl regenerating its certificate for the new
name. Host tests: 357 forgetest unit tests, forgectrl clean under
-Werror, tls_test and sanitize_test.
2026-09-10 07:18:27 -04:00
ScottW514 d858a23f45 Carry the panel stylesheet into forgetest
The acceptance page shares theme.css with the panel byte for byte, and forgectrl's copy gained the setup header's Download logs button. CI caught the drift, which is what that check is for; the pinned revision decides, so the copy follows.
2026-09-09 19:53:09 -04:00
ScottW514 c7b80ab2e8 A test that does not hand the machine back fails
The baseline has always examined the machine after every run and recorded
what the run left behind. It did nothing else with it: the leftovers went
to the log and the evidence, and the test still reported PASS. So a check
could measure correctly, walk away with the machine in a state nobody
chose, and be recorded green.

That is how the purge fan came to be left off by the airflow check. The
leftover was not even watched, but had it been, it would have been noted
and the test would have passed anyway, and an operator would still have
met the airflow hold at their first fire.

A post-run leftover now fails the run. One the baseline put back fails it
too: the restore is the bench cleaning up after a defect, not the defect's
absence. The message names what was left.

The baseline watches the head as well as the motion side now: purge air
on, which is how the machine idles, and the lens motor at its hold current
in half step, which the lens checks and the sheet cards take and must hand
back. The airflow check proves the machine is whole rather than merely
measured: afterward the purge fan must read commanded-on and must draw
above the floor the check itself just wrote.

The pin takes forgectrl 0.1.13 (a2d73ef), which restores the idle posture
after a diagnostic, fixes the lens session's takeover flag, and gives the
setup a Download logs button, since the panel's Logs tab is unreachable
until the setup is complete.

Expect this to find things. A test that has been handing the machine back
imperfectly has been passing until now, and the first campaign under the
rule is where that shows.
2026-09-09 19:44:24 -04:00
ScottW514 468e92655a Teach the manifest guard test about the version file
The guard keeps scripts/manifest-from-tree.py and forgefirm-image-manifest.bbclass computing the same identity, so it pins the skip list by value and the new suffix failed it. It now expects the version file in both places, and a second case proves the point the change exists for: setting the release number leaves the layer hash where it was. My miss: I proved the identity by hand and pushed without running the unit suite first.
2026-09-09 18:28:47 -04:00
ScottW514 d898b5659d Give the release version its own file, outside the layer content hash
Setting the release number was a platform change. FORGEFIRM_RELEASE sat
in forgefirm-image.bb, the recipe hashes as content of meta-forgefirm,
and a change to the content of a layer invalidates every acceptance
result. So a version bump threw away the campaign that was meant to
authorize that very release, and the number therefore had to be decided
before the image the campaign ran on. Nothing said so: the release-flow
page went straight from the kas configuration to the artifact and the
pipeline, while the gate quietly required the recipe value, the rootfs
stamp, the archive's meta-version and the tag to agree. v0.0.1 was cut
on a tree whose number happened to be right; the next one would have
cost a second campaign to discover the rule.

The number moves to forgefirm-release.inc, which carries it and nothing
else, and the manifest leaves that file out of the layer content hash
exactly as it leaves out the component pin files
(FORGEFIRM_MANIFEST_VERSION_SUFFIX, and the same list in
scripts/manifest-from-tree.py, which computes the identity on a
workstation and must agree byte for byte). release.sh reads the number
from the new file.

The version is metadata, not platform content, and this only makes the
manifest say what it already meant: the version string was already
outside the identity hash, and it was the file carrying it that defeated
that. Nothing is weakened. release.sh still requires the number to equal
the rootfs stamp, the .fw meta-version and the release tag, and
image.health still compares the stamp on the running machine with the
manifest's.

Proven: the tree manifest is byte-identical across a bump from 0.0.1 to
0.0.2 (identity a64e51b8e5ecca0af683d4f0 either way, the meta-forgefirm
layer hash unchanged), where before the two differed. bitbake resolves
FORGEFIRM_RELEASE=0.0.1 and FORGEFIRM_VERSION_STRING=v0.0.1 for the
release image through the new require, and the dev image still overrides
the string with its build timestamp.
2026-09-09 18:12:03 -04:00
ScottW514 6967308485 Release v0.0.1: the acceptance artifact for this tree
Campaign c-20260909204732-3fa9 on image 20260909193551, 83 tests, 83 satisfied (62 inherited), release authorized. It replaces the artifact of the release that was withdrawn: that one authorized an earlier rootfs, and the gate compares against the rootfs it is asked to sign.
2026-09-09 17:55:33 -04:00
ScottW514 64301d3221 Keep the configuration files inside /data/forgefirm; pin the three components
ForgeFIRM's own files live under /data/forgefirm; two configuration
files did not. The machine settings sat at /data/forgefirm.conf, in the
root of /data beside the factory's own files, and the cloud-mode
configuration sat at /data/etc/gfhome.conf, inside a directory the
factory owns. Both move:

  /data/forgefirm.conf   -> /data/forgefirm/forgefirm.conf
  /data/etc/gfhome.conf  -> /data/forgefirm/gfhome.conf

There is no migration: only the bench has ever run this firmware.
/data/etc now holds only the factory's wpa_supplicant.conf.

The acceptance check of the file modes reads the settings file at its
new path, and the two bench tools that read it directly follow. The
pins move to the revisions that carry the change, forgectrl also
bringing the fix that reads the module's disabled state as idle:

  forgectrl           468ee21 (0.1.12)
  grblhal-glowforge   9ee624b (0.1.10)
  forgefirm-app       56f134a (0.1.28+git)

The lock moves meta-openglow to b7ad6d9, which pins python3-gfhardware
on the same revision. The four upstream layers stay where they were:
`kas lock --update` moves every floating repository, and a release is
not the place to take poky, meta-openembedded and meta-freescale along
for the ride.
2026-09-09 15:32:54 -04:00
ScottW514 61030da69b Pin forgectrl on the disabled-is-idle fix
A machine out of the box sat in the module's power-on state, disabled, for the whole of its first run, and every idle gate read that as busy: the setup's sensors check refused to start, settings writes answered 409, and the cooling engine held cooldown airflow from boot. forgectrl now reads disabled as idle (the state means no program in progress), with fault and underrun still busy until acknowledged and an unreadable state still failing closed.
2026-09-09 14:42:00 -04:00
ScottW514 0107c2cad4 Remove the cold-build workflow: Yocto builds on the build host only
The yocto-cold-build workflow and its kas/ci.yml overlay built the release image on a hosted runner as a reproducibility probe. It never ran to completion, its first dispatch (2026-09-09) stopped on the runner's user-namespace rule, and a probe nobody runs is a trap. Every Yocto build, the release included, runs on the build host; the release proof is the local pipeline (release.sh) and the bench campaign. The pre-publish checklist loses its self-containment line to match.
2026-09-09 13:28:05 -04:00
ScottW514 867b1938e4 Cold build: allow unprivileged user namespaces on the noble runner
BitBake isolates the network of its tasks with a user namespace, and the ubuntu-24.04 hosted runner's AppArmor profile refuses that to an unprivileged process, so the cold build stopped before its first task (run 34381825302). The workflow lifts the restriction for the run; nothing in the layers or the image changes.
2026-09-09 13:22:12 -04:00
ScottW514 b2f50ad765 Release v0.0.1: pin meta-openglow, refresh the lock, add the acceptance artifact
The kas configuration takes the pinned-remote meta-openglow block, with
its commit in the lock file (d655e1e, the read-only rootfs), so a fresh
clone builds the release without a sibling checkout. The lock keeps the
upstream layers where they were.

releases/v0.0.1 carries the acceptance artifact the bench exported for
this image: campaign c-20260909160235-7649 on 20260909150456, 83 tests,
83 satisfied, none inherited, release authorized. The release gate
recomputes every test's fingerprint from the manifest inside the release
rootfs and signs only when the recorded results agree.
2026-09-09 13:15:37 -04:00
ScottW514 4a95595afd Pin forgectrl on the gate that waits for the enclosure; add its test
forgectrl now holds the motion check while a lid or the interlock is
open instead of starting the controller unverified: GET /mode reports
controller "waiting" with why, the button blinks amber, and the check
runs when the enclosure closes. motion.gate-waits-for-lid drives the
fixture's lid channel: the lid opens, forgectrl restarts, /mode must
read waiting with why naming the lid, no pid, motion unverified, the
button amber (sampled over a blink period: the smooth trigger's target
reads 0 through the off half) and no probe line in the log; the lid
closes, and the controller must come up verified with MOTION OK on the
first probe.

Proven on the bench reference: PASS, the controller verified 6.5 s
after the lid closed.
2026-09-09 11:03:33 -04:00
ScottW514 8af8b197ee Mount the rootfs read-only on both images
The rootfs mounted read-write, so a slot ran with its own files open to
change, and the factory-slot mounts rode along on the release image.
Both images now carry the read-only-rootfs feature: the ro root line and
the rcS default, the volatile links made at rootfs time, a writable copy
of /var/lib at boot, a build failure for a post-install that needs the
machine, and the removal of shadow, base-passwd, update-rc.d and
update-alternatives.

What must last or change at run time is handled file by file:

- forgefirm-users renders the four account files from the record into
  /run/forgefirm/accounts and bind-mounts each copy over its /etc file
  (useradd and the rest are gone with shadow); a render writes through
  the mount, and the image's own files apply until the first render.
- forgefirm-banner bind-mounts a copy of /etc/issue and writes the
  address block through it.
- sshd keeps its host keys under /data/forgefirm/ssh, so the fingerprint
  survives updates; both sshd configs carry the same HostKey lines.
- forgefirm-logging passes logrotate a state file under /var/run
  (logrotate refuses to run without one).
- forgefirm-persist points the boot timestamp and the random seed at
  /data/forgefirm.

The dev image appends the /factory slot mounts, without nofail (busybox
mount hands it to the kernel, which rejects it). The rootfs command
entries lose their semicolons: on scarthgap the value is the task's
vardeps, split on whitespace, so "name;" left the function body out of
the signature and a changed body did not remake the rootfs; with the
bodies tracked, the dev image's DATETIME string needs a vardepsexclude.
release.sh gains the read-only gate (root ro, no /factory line,
ROOTFS_READ_ONLY=yes, host keys on /data). image.health checks the
mounts, the account binds, the banner bind, the host keys and the
dev-only /factory mounts.

Proven on the bench reference (dev image 20260909140901): / ro, /data
rw, /var/lib a tmpfs copy, the four account files and /etc/issue bound
from tmpfs, the host keys in /data/forgefirm/ssh, no "Read-only file
system" line in any log; forgectrl.auth and commission.account-login (a
temporary account rendered, logged in over HTTPS and removed again),
kernel.latch-locked-idle and motion.liveness-probe PASS; logrotate runs
with the volatile state. forgetest unit tests 335 OK; both images build
clean, and debugfs on the built rootfs shows every setting above.
2026-09-09 11:02:10 -04:00
ScottW514 2936890eaa Pin forgectrl on the revision CI passed
forgectrl 6040e64 carries the three fixes behind the pinned 93fb22e: the
test header reached the way the sibling tests reach theirs, the key added
to /status carried in the panel dev-server mock, and the lens test's
carriage driven by the sweep's own steps rather than by a clock. The
shipping behavior is unchanged from 93fb22e; the pin names what passed.

Verified with bitbake -c fetch.
2026-09-09 08:43:45 -04:00
ScottW514 1ea686ef85 Re-pin forgectrl and grblHAL-glowforge on the fixed revisions
forgectrl 93fb22e fits the lens outcome text in the buffer the supervisor
shares with the probe. grblHAL-glowforge 97be92b takes the lens reference
from the realtime hook rather than settings-changed, so the controller no
longer overwrites the Z it just referenced.

Both verified with bitbake -c fetch at these revisions.
2026-09-09 08:28:39 -04:00
ScottW514 093a7bbbde Stage the lens reference in the laser-stream harness
The Z session used to reference the lens itself with M103, which is gone:
the daemon sweeps the lens onto its hall edge before a controller starts
and leaves a marker, and the controller opens the Z envelope on that. No
daemon runs behind the harness, so nothing wrote the marker and every Z
move was refused, which is what the session's first move ran into.

The runner now writes the marker the daemon writes, and the session pins
lens_hall_edge_z_mm so the moves are counted from a known height: Z3 is 9
half-steps on the screw and Z4 is 12, so a 1 mm move up and back is 3
steps each way.

This is also the check that caught the controller overwriting its own
referenced Z at start, which the panel could not show.
2026-09-09 08:27:32 -04:00
ScottW514 3e5c1a342a Pin forgectrl and grblHAL-glowforge 0.1.9: the lens reference at start
forgectrl dd40dc8 takes the lens onto its hall edge in the motion-verify
window and gates the spawn when it cannot, and carries the per-axis
anchor, homed_axes on /status, and the panel reading Z from its own bit.
grblHAL-glowforge ef0f764 takes that reference as it loads its settings,
re-zeroes the kernel counters the daemon's GPIO steps never reached, and
drops M103.

Both verified with bitbake -c fetch at these revisions.
2026-09-09 08:08:47 -04:00
ScottW514 5f19ae01cd Follow the lens reference into the harness and the catalog
The lens now takes its hall-edge reference before any controller starts,
so Z is referenced on every start and M103 is gone. The laser-stream
harness opened its Z session by referencing the lens the way a
commissioning card did; it no longer has to, because Z is already open by
the time the session runs.

forgectrl.panel-serves gains the assertions for the per-axis reference:
homed_axes is an axis mask, homed agrees with it, and with a controller
running Z is referenced and reads inside the lens reach the same document
reports. That last check is the one that catches a panel showing nothing
for a Z the controller holds.

commission.check-motion already exercised the new path, because the
motion wizard's probe runs the same sequence the supervisor does, so its
covers map gains lenshome.c and its description names the lens reference
and the hard fault behind it.
2026-09-09 08:06:04 -04:00
ScottW514 74fa8a4ee6 Update README 2026-09-09 07:01:15 -04:00
ScottW514 be9e4198ba Assignment 2026-09-08 17:06:34 -04:00
ScottW514 291f75815b Assignment 2026-09-08 16:48:43 -04:00
ScottW514 c2ca2686a7 Assignment 2026-09-08 16:21:37 -04:00
ScottW514 722de73a39 pins: forgectrl 0.1.7, grblhal-glowforge 0.1.7, forgefirm-app 0.1.26+git
The three components carry documentation-only commits: the README
becomes an index card, AGENTS.md lands, and the comments that pointed at
the retired BRINGUP.md now point at the documentation site.

The pins move because the tree is already past the image that
20260907224817 built: a comment in the kernel config fragment
glowforge.cfg changed the meta-glowforge-bsp layer content hash, which
the manifest counts as a platform change. Measured against that image's
manifest: meta-forgefirm and meta-openglow-core are unchanged,
meta-glowforge-bsp is not. A fresh image build and a full acceptance
campaign therefore precede any release, and holding these pins back
would buy nothing.

Verified: every pinned revision is on its public remote, and
bitbake -c fetch resolves all of them.
2026-09-08 13:37:13 -04:00
ScottW514 e36a322ec5 Retire BRINGUP.md and CAMPAIGN-LOG.md
Every fact in the two documents is now on the documentation site, which
is the single source of truth. This repository carries no project
documentation any more: it is the build and release base plus the
acceptance tool, the bench tools and the fixture firmware.

BRINGUP.md was the runbook, the hardware facts bank and the open-work
list. CAMPAIGN-LOG.md was the dated record of how each result was
obtained. What replaces them: the site for present state, and the
commit message for the record of what a change did and how it was
proven, so the change and its record stay together. Local open work is
the developer's own file at the tree root and is not tracked here.

README.md becomes an index card: what this is, build, test, and where
the documentation is.

The release pipeline tags the documentation. Firmware on a machine
needs the documentation that agrees with it, so release.sh now tags the
forgefirm-docs checkout with the same v<version> as the release, and
prints the command that pushes the tag with the release. The checkout
must exist and be clean, which is a new gate before the signature.
FORGEFIRM_DOCS_DIR names the checkout (default: the sibling one) and
FORGEFIRM_DOCS_SKIP releases without a tag, loudly, and is never the
default. The tag is made at staging and pushed with the release, never
before: a documentation tag for a release that never shipped is worse
than no tag.

No catalog consequence. release.sh is host-side and is in no image.
The commission.py change is one sentence of a test description, not
behavior. accel_crash_probe.py and the kas header lose pointers to the
retired files.

Checks: bash -n and sh -n on release.sh, and the tracked trees carry no
reference to either retired file.
2026-09-08 13:29:43 -04:00
ScottW514 81df43e7a4 pins: grblHAL-glowforge 0.1.6, the planner buffer index fix
grblhal-glowforge e2ba043 carries the core fork at 362577d, whose block
ring index covers $398 up to 1000. Fetch-verified with bitbake -c fetch.
Pin bump only.
2026-09-07 18:40:44 -04:00
ScottW514 59c4516c00 The planner buffer depth harness; the XY microstep set is in
scripts/bench/planner_blocks_test.py restarts the null-sink controller
at $398=400 and at 1000 on one settings store and requires an answer on
the port, the depth in the status report and a move to Idle. It runs in
the grblHAL repo's CI; registered on the bench page and in the README.
BRINGUP: $398 runs over its whole range, the spin item is closed, and
the XY microstep item is committed, pushed and pinned. CAMPAIGN-LOG: the
landing and the index fix, with the host and bench proof.

No catalog consequence beyond the core submodule the motion tests
already cover; the coverage lint is clean.
2026-09-07 18:38:39 -04:00
ScottW514 0dc8c746e5 pins: forgectrl, grblHAL-glowforge and the app at the XY microstep revisions
forgectrl 0.1.6 (b1eee4d): the xy_microsteps setting and the quiet hold route.
grblhal-glowforge 0.1.5 (48d5f1d): the XY scale, tick and stop ramp derived
from the microstep mode, plus the producer lead ceiling.
forgefirm-app 0.1.25+git (d1c47b8): the checked header keys note; the
python3-gfhardware pin in meta-glowforge-bsp moves with it.

Fetch-verified with bitbake -c fetch. Pin bump only: the acceptance
consequence is the catalog invalidation the manifest already derives.
2026-09-07 18:11:18 -04:00
ScottW514 8fc5250d6d XY microstep modes: the baseline, the catalog test and the bench tools
The baseline derives x/y_mode, step_freq, ramp_rate and the configured
markers from the xy_microsteps setting; ramp_rate joins the GRBL
controller's set (the driver writes it; the cloud client runs at the
module's). motion.microstep-modes cycles 8, 16 and 32: the save restarts
the idle controller, the kernel reads the mode with its tick and ramp,
$100/$101 are the mode's and a typed $100 is overwritten, a 40 mm jog at
top speed returns to Idle with the kernel counters over the mode agreeing
with the commanded travel and the accelerometer seeing the head move; the
setting is put back as found.

Bench tools: xy_mode_test.py (the null-sink harness the grblHAL CI runs),
raster_dry.py (a top-speed raster per mode), xy_pattern_accel.py (the
operator's pattern from home with the machine silent and the head
accelerometer listening), arc_tolerance_sweep.py (a $12 ladder on the 9
in circle: the chord rate the protocol loop feeds, about 300 a second, is
the ceiling, not the core), and the xymode and xycircle live drills.
BRINGUP carries the present state and the facts; CAMPAIGN-LOG the dated
record, including the planner-blocks spin (a $398 of 255 or more loops
forever at start, a core bug) and its recovery.
2026-09-07 18:05:28 -04:00
ScottW514 bfb5cb27d8 The tail after Idle is one queue depth, and it does not grow
The open item "Idle before the kernel drains" asked to decide between
holding Idle until the kernel drains and stopping the continuation pads
growing the lag. Measured first: on the machine, four chained 50 mm jogs
against cnc/state give 171, 175, 177 and 176 ms after Idle. One queue
depth, flat. The growth the item described is gone, so neither driver
change was made and the item is closed.

Host side agrees and says where the mechanism lives. Stream bytes are the
time axis, so a dumped stream's length is how long the machine plays it:
chained jogs produce 35755 bytes each with no growth, and the churn
session holds at 64790 bytes at a producer lead of 2 or 10 ms. Only above
the lead ceiling does it inflate. Rule 17 now holds that stream to a
budget derived from the job rather than to a recorded number, so the
inflation regime cannot return unnoticed; it fails at 5899 and 8005 ms
with the ceiling lifted and passes at 2301 against 3800.

BRINGUP carried a second error. It said every forgectrl path that stops
the controller after motion waits for cnc/state idle. super.c says
outright that POST /controller/stop is not idle-gated, because it is also
the emergency lever, and safes the machine with cnc/stop and the latch
before the signal instead. The mode switch, the cooling gate and the
daemon shutdown do gate on machine_is_idle(). Both the tail figure and
the gating claim are corrected, and CAMPAIGN-LOG carries the measurements
and the decision not to hold Idle.

Acceptance: rule 17 is a host harness rule in the grblHAL repo's CI, not
a catalog case. The catalog is unchanged because no machine behavior is:
the driver change is a bound on an out-of-range knob.
2026-09-07 12:19:16 -04:00
ScottW514 527d864eaf Register the Z envelope harness and record the rebase
A new scripts/bench python file must be entered in the bench registry as
well as in that directory's README: the registry is what the acceptance
page reads, and a file in neither is a tool nobody can find. The harness
joins the other two host-side CI harnesses, marked not a bench-page tool
for the same reason they are - it drives the host-built null-sink
controller, not the machine.

CAMPAIGN-LOG gets the dated record of the rebase onto core build
20260905: the settings-struct measurement taken before the machine was
touched, the harness and the two broken builds it was validated against,
and the bench results.
2026-09-07 11:19:04 -04:00
ScottW514 e6e5ca1c94 pins: forgectrl e27b412 (0.1.5), grblhal-glowforge 96906b6 (0.1.4)
forgectrl carries the GRBL settings store into the data directory, which
the acceptance suite already asserts. grblhal-glowforge follows the core
off the deprecated settings-changed event and moves the core submodule to
the fork rebased onto build 20260905.

Both revisions fetch.
2026-09-07 11:08:58 -04:00
ScottW514 79d2c07734 The Z envelope survives a settings write
The Z soft limit belongs to the driver, not to $20. glowforge_homing.c
owns sys.work_envelope, sys.homed and sys.soft_limits for Z, because Z is
always referenced, to the lens hall edge or to where the lens stands, and
the core knows neither. The core recomputes both masks from the settings
and drops Z when it does: $20 clears the soft-limit mask inside its
setter, and a $13x write clears the homed bit for the axis as well.

z_envelope_test.py drives the null-sink controller over TCP and holds the
rule. It checks that an unreferenced Z is collapsed to where the lens
stands and refuses a move each way, that X and Y stay free so the
reassert is Z's alone, and that neither write frees Z. It restores $132.

The harness runs in the grblHAL repo's CI, next to the laser stream and
lifecycle harnesses.
2026-09-07 11:02:54 -04:00
ScottW514 b0fa4ccaf5 A release publishes the source of the software it installs
The release build merges kas/source-bundle.yml, which turns on the Yocto
archiver: the upstream source of each recipe as upstream publishes it, the
patches with their series file, and the recipe with its includes. The
overlay adds tasks only, so the image manifest is unchanged and an
acceptance result still applies; proven on the build host, where the
archiver build and a plain rebuild of the same tree give the same
content_sha256.

scripts/source-bundle.py packs forgefirm-source-v<version>.tar.gz: the
archives, both license manifests, the license texts, the ForgeFIRM layers,
the kas configuration, the layer revisions and the build identity of the
image. What the bundle must hold comes from the image, not from a list in
the script: every recipe of license.manifest and image_license.manifest
whose license is in the include list must have an archive, or the release
stops with the recipe named. release.sh attaches the bundle and covers it
with sha256sums.txt; FORGEFIRM_SOURCE_SKIP=1 bypasses deliberately.

On the build host: 68 of the image's 111 recipes carry source, 313.7 MiB,
under the 2 GiB limit of a release asset.

No acceptance catalog consequence: the change is release tooling on the
build host and puts no file and no behavior on the machine. The host-side
proof is forgetest/tests/test_source_bundle.py, which holds the license
decision, the choice of archive and the refusal.
2026-09-07 09:49:18 -04:00
ScottW514 9f715f08de The GRBL settings store lives in /data/forgefirm
BRINGUP names the new path in the standalone start line and the
stored-settings note. forgectrl.panel-serves checks, in GRBL mode, that
the store is /data/forgefirm/EEPROM-glowforge.DAT and that nothing of it
remains at the top of /data. super.c is already in its covers map.
2026-09-06 22:22:33 -04:00
ScottW514 a86d66049f forgetest: the invalidate-all notice ends when a campaign starts after it; its epoch stays 2026-09-06 22:03:56 -04:00
ScottW514 bd624410a2 docs: the full campaign passes on image 20260907005922; BRINGUP item 8, the first-run commissioning, is closed 2026-09-06 21:58:30 -04:00
ScottW514 143ef11a60 forgetest: the liveness test waits for the probe's log line; forgectrl pinned at 0.1.4 (the probe's detail text); the campaign log records the first campaign 2026-09-06 20:57:56 -04:00
ScottW514 0df0162c8d forgetest: the sheet test covers the runner's header; the hollow generator entry is dropped 2026-09-06 20:23:03 -04:00
ScottW514 54f65746cd forgetest: the page's theme follows forgectrl's at the pinned revision 2026-09-06 20:18:47 -04:00
ScottW514 60e06c77f0 forgetest: the lens stall drills are named as shell-only in the bench registry; the campaign log records the pushes and the build 2026-09-06 20:09:44 -04:00
ScottW514 8795a6d527 pins: forgectrl, grblHAL-glowforge, and forgefirm-app at the commissioning revisions 2026-09-06 19:56:29 -04:00
ScottW514 97287aa6a9 commissioning: the layer, the acceptance tests, the harness rule, the docs, and the bench drills
meta-forgefirm: the forgefirm-users init replays the account at boot;
sshd refuses root and empty passwords and runs only while the panel
turns it on; the release image keeps an empty root password for the
console; the console banner; avahi announces forgefirm.local; https in
libmicrohttpd and ulfius; the panel on 80 and 443; the license bundle on
the rootfs; release.sh checks the root policy on the built rootfs.

forgetest: the commission suites (commission, commission_dark,
commission_sheet: 23 cases); the runner turns cloud mode on with the
typed phrase for a test that declares it; the baseline's motor_lock is
0; the log-export test checks the bundle for the camera key; the record
helpers write bytes as given and join the daemon's paths as POSIX. The
stream harness gains rule 24: a hold verdict is held again after a
resume. Bench drills: lens_travel.py and lens_stop_accel.py.

Docs: BRINGUP carries the present state; CAMPAIGN-LOG carries the dated
record.
2026-09-06 19:56:05 -04:00
ScottW514 ff9796cde6 docs: the 2026-09-01 audit is retired
Its last finding, the request-body cap, ran on image 20260904131106 in
campaign c-20260904132654-d731. forgectrl.auth carried the case: a 4 MiB
body from a client with no token was refused with 403, and the daemon
answered /status in the same second. The campaign is 56 of 56,
authorized.

That was the only audit item left in Next work, so item 9 goes and the
list is items 1 to 8. The record of the campaign, and of the retirement,
is in CAMPAIGN-LOG. The audit file is deleted, as the 2026-07-03 and
2026-08-13 audits were before it.
2026-09-04 10:15:10 -04:00
ScottW514 da0299924f pins: forgectrl at the revision the campaign ran
The request-body cap is on the bench. Image 20260904131106 carried
forgectrl ff89288 through a build overlay, and campaign
c-20260904132654-d731 authorized it, 56 of 56.

The pin now says what the campaign ran.
2026-09-04 10:15:02 -04:00
ScottW514 637fb67007 docs: the audit's follow-through closes, its last finding stands alone
Three items in Next work tracked the 2026-09-01 audit: the remediation
follow-through, the PIC readings, and the deferred six. All three are
done, on one image, and the campaign the release gate asks for passed on
it, so they belong in the record rather than the open list.

What remains of the audit is one finding, the request-body cap, which is
now fixed and host-proven and needs an image and a campaign. It takes
their place as item 9.

forgectrl.auth gains the case that guards it: an oversized body from an
unauthenticated client is refused and the daemon is still serving after.
2026-09-03 18:41:46 -04:00
ScottW514 9c6b13e510 pins: the revisions the authorized campaign ran
forgectrl 8ef9509 at 0.1.2, grblhal-glowforge 2f5edee at 0.1.2, and the
forgefirm-app recipes at dd0ebf3, 0.1.23+git. These are the sources the
full campaign on image 20260903211413 passed against, now that their
repositories are pushed and the build no longer needs the local overlay.
2026-09-03 18:12:01 -04:00
ScottW514 5a74213919 docs: the full campaign on 20260903211413, 56 of 56, authorized
The campaign the release gate asks for, run on one flash with nothing
inherited and nothing hot-deployed. The entry records how the pause test
passed, not only that it did: its emission trail drains to zero and stays
there, against the earlier run's second kernel run inside the pause, and
the difference between them is that the actuator made the presses.

Also recorded: the two harness faults found on the way, the actuator
dropping off the network twice and being asked for silently, and what
now goes in the record so a repeat is answerable.

BRINGUP moves to the campaign image and drops the campaign from the owed
list, leaving the push and the pin bumps.
2026-09-03 18:09:05 -04:00
ScottW514 d6f648b539 forgetest: presence is proved once per test, not once per ready gate
The ready gate lives inside the arm-and-fire helper, and the kill drill
calls that helper twice, once for the expected stop and once for the
SIGKILL. So the presence gate asked the operator for a second press part
way through a test they had already proved themselves present for, with
the actuator standing by holding the presses. It is the only test in the
catalog with two ready gates.

The second gate now returns at once. Its setup line still goes up,
because the second half may want the scrap moved, but there is no press
to make.
2026-09-03 18:08:16 -04:00
ScottW514 2f330715f3 forgetest: record the actuator's radio signal and uptime with every run
The bench actuator is an ESP32 on wifi and it reports its own signal
strength, which reads -83 dBm here, close to where an association
starts dropping. It has vanished twice tonight and taken a live test
with it, and neither time did the record say anything a reader could
use: only that it was gone.

Both numbers now go into the run's evidence, so the next drop says
whether the link faded or the box restarted.
2026-09-03 17:08:02 -04:00
ScottW514 ba0bf41749 forgetest: the operator proves presence at the machine, the bench presses
A live test asked a person to click Ready on a page and then make
timing-critical presses in the middle of a burning cut. That is how
tonight's pause test became unanswerable: the actuator had dropped off
the network, the harness fell back to the operator without saying so,
and afterwards nobody could tell a second press from the machine
resuming on its own.

Where an actuator is up and wired to the button, the ready gate now
takes a press on the machine's own button as the presence check, and the
actuator performs every press in that test. The button does nothing at
Idle, so the press is only a presence check, and the gate waits for the
release so it is never read as the arm press. With no actuator the
operator does the presses and answers on the page, as before.

An actuator lost after that takeover is now said out loud, in the log
and in the evidence, instead of quietly becoming a person's press.

The live-fire cue was four lines of machine-shaped prose. It is now what
a person needs: protection, exhaust, extinguisher, scrap, lid.
2026-09-03 16:55:34 -04:00
ScottW514 da0c70703d docs: the shakedown on 20260903163543, and the arm fix under live fire
The whole catalog ran green, 56 of 56, but on a hot-deployed suite file,
so it authorizes nothing; the entry says so first. What it did was find
the one real defect, the pic-soc-load bound that decided on a count of
noise, and put the arm-acknowledgment fix under live fire: the emission
witness passed with its airflow witness finding nothing, on the same
instant button press that produced last night's burn.

Also recorded: the boot clock that never stepped, and the bench fixture
dropping off the network mid-queue with no address pinned to fall back
to.
2026-09-03 15:16:24 -04:00
ScottW514 7968ebd347 kernel.pic-soc-load: the move bound is a third of the split, not half
The check decided on one count of noise. Measured over eleven runs on the
bench the settled reader moves 3 counts off the idle regime (once 2, once
4) against a control split of 6 (once 7), so a half-split bound sits
exactly on the median: two of those eleven runs failed while the machine
read identically to the nine that passed, and the queue stopped on one of
them.

The kernel's spin is its own load level, a couple of counts under a
Python spin, so a settled reader lands above the idle regime without
reaching the busy one. Asking it to reach halfway was asking for
something the mechanism does not promise. What the check has to catch is
a settle that overshoots and lets the conversion fall back to idle, which
reads as no move at all, and a third of the split catches that with a
count of margin either way. The split collapsing is still the primary
proof, unchanged above.
2026-09-03 14:28:48 -04:00
ScottW514 0845ab32a1 docs: build p33, image 20260903163543, the campaign's image
The record of what the first attended run really found and what it cost:
the verdict had no run-session identity, so the arm opened its window on
the verdict computed for the idle session before it. The build entry
lists the four components taken from local commits and the checks made
on the built images.

BRINGUP moves to the new campaign image and says why nothing inherits.
2026-09-03 12:46:39 -04:00
ScottW514 db9acf9910 forgetest: witness the airflow behind the beam; harnesses carry the armed flag
The cooling verdict now carries the engine's own armed flag, so the
stand-in engines in both null-sink harnesses publish it. The lifecycle
harness gains two cases: an engine that never takes the armed window
must produce a refused arm and no emission, and one that takes it a
couple of seconds late must produce a wait and then a normal arm. The
late case is the one that proves the controller keeps reading the
verdict while it is blocked in the arm; without that every job would
fail there.

The emission witness gains the bench form of the same rule: no sample
may show the laser firing while the cooling engine reports a phase that
runs the fans at their idle duty. That is what a burn with no airflow
looks like from the outside, and nothing in the catalog looked for it.
2026-09-03 12:26:59 -04:00
ScottW514 aa0d86635c CAMPAIGN-LOG: the attended set stopped at its first test
The record of the live-fire failure on 2026-09-03: a job fired for about a
second with idle airflow under a warm-up hold that arrived one tick later,
the leftover start-gate setting behind it, the engine's arm-time gate gap it
exposed, and what is owed before any further live fire.
2026-09-02 21:59:31 -04:00
ScottW514 ec9e9860c9 CAMPAIGN-LOG: the unattended set on image 20260903011655, 44 of 44
The record of the campaign's unattended set on p32, the numbers the two new
kernel drills and the air-assist offset diagnostic gave under the settle,
and the forgetest restart that interrupted the last test, with the rule
that keeps it from recurring.
2026-09-02 21:49:18 -04:00
ScottW514 9e862971cf kernel.pic-soc-load: the settle's proof is the move off the idle regime
The kernel's spin is its own load level, a count or two under a Python
spin, so the third check no longer compares the settled level with the
Python-spin control; it requires the settled idle reader to have moved at
least half the control's split off the idle regime. Bench: PASS on image
20260903011655 (control split +6, settled split +1).
2026-09-02 21:46:35 -04:00
ScottW514 cf7490de78 docs: build p32, image 20260903011655, the campaign's image
BRINGUP item 9: items 10 and 11 are on image 20260903011655; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p32 build and its checks.
2026-09-02 21:22:35 -04:00
ScottW514 92c69f23fa kernel.pic-soc-load drill; the PIC worked backward from its firmware
The PIC16F1713 converts its inputs in a free-running loop (10 channels,
about 0.30 ms a loop) and a read returns the last conversion of that
channel; the count follows the SoC's load at conversion time (idle 659,
busy 665 on the coolant thermistors, both tight; every channel shifts in
proportion to its count; the step lands within one PIC loop of the CPU
changing state, with a regulator's overshoot in each direction). The
kernel.pic-soc-load drill replaces kernel.pic-pacing: 200 reads after 3 ms
of sleep and 200 after 3 ms of spinning, with the module's settle off
(the control, reported) and on (the claim: the two agree). The catalog
counts 56 tests, 0 uncovered.

BRINGUP item 10 and the facts bullet describe the mechanism and the fix;
the CAMPAIGN-LOG entries record the first pass of the campaign, the PIC
study, and the mechanism's proof.
2026-09-02 21:15:34 -04:00
ScottW514 7c45642f05 image.health: the watchdog's proof is WDOG1's own WCR, not the sysfs state
The sysfs state of watchdog0 says whether a process holds the device, and
none does: the kernel's core feeds the boot-armed hardware. The check
reads WCR through /dev/mem (WDE set, a 60 s period) and expects the state
to read inactive. Bench: PASS on image 20260903003213 with WCR 0x771f. The
BRINGUP facts bullet says the same.
2026-09-02 20:46:48 -04:00
ScottW514 5d2aa41a1c docs: build p31, image 20260903003213, the campaign's image
BRINGUP item 9: items 10 and 11 are on image 20260903003213; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p31 build and its checks.
2026-09-02 20:38:45 -04:00
ScottW514 358c287891 kernel.pic-pacing drill; BRINGUP item 10 and the record
kernel.pic-pacing reads a coolant thermistor twice back to back, 300
pairs, with the module's pacing off (the control, reported) and on (the
claim: the second read agrees with the first). The bench proof for the
module's pic_gap_us pacing; the catalog counts 56 tests, 0 uncovered.

BRINGUP item 10 and the facts bullet describe the pacing as it is; the
CAMPAIGN-LOG entry records the change and its proof.
2026-09-02 20:30:41 -04:00
ScottW514 1f2c795d8c docs: build p30, image 20260903000529, the campaign's image
BRINGUP item 9: the deferred batch is on image 20260903000529; owed is the
flash, the fresh-boot baseline, and the full campaign, then the push and
the pin bumps. CAMPAIGN-LOG: the p30 build and its checks.
2026-09-02 20:13:31 -04:00
ScottW514 9711a7fa06 The audit's deferred findings: the checksum exclude, the resume-lead drill, the record
B-16: BB_SIGNATURE_LOCAL_DIRS_EXCLUDE in the distro conf names __pycache__
and .pytest_cache, so a workstation's bytecode caches never enter a
file:// checksum (proven in the build VM: a cache under the package leaves
the fetch task alone, a source change reruns it).

kernel.resume-lead: two phases behind one takeover at a 1 kHz tick. E, a
resume whose lead is longer than the data ends at end-of-data within a few
ticks (a lost end-of-data would show as 255 ms). L, a 1000-byte lead over
FIRE bits with the latch unlocked and the chain unarmed keeps the FIRE
line low through the lead and drives it from the waypoint byte on. The
bench proof for the module's K-4 and K-8 fixes; the catalog counts 55
tests, 0 uncovered.

BRINGUP items 9 and 11 and the CAMPAIGN-LOG entry record the batch and the
campaign rule: no campaign until every audit finding is on one image.
2026-09-02 20:03:22 -04:00
ScottW514 abccbf48e5 docs: build p29 on the pushed pins, image 20260902230436
CAMPAIGN-LOG: the push in CI order, the pin bumps, and build p29 with its
checks. BRINGUP item 9: the remediation is pushed, pinned, and built; owed
now is the flash, the fresh-boot baseline, and the full campaign on the
image, which the release gate asks for.
2026-09-02 19:15:52 -04:00
ScottW514 77c11c52e4 forgetest CI: run the python steps with -B so no bytecode enters the recipe fetch
B-16: the forgetest recipe fetches the package directory whole, and the
file fetcher's checksum is taken before do_unpack drops __pycache__, so a
host `__pycache__` written by running the tests moves the recipe's task
hash with no source change. The three python steps run with -B now;
`src-sync` to the build VM excludes the caches as well. A developer who
runs the tests locally without -B still creates them; there is no exclude
on the file fetcher, so this is the floor, not a cure (BRINGUP item 11).
2026-09-02 19:09:07 -04:00
ScottW514 05cf7341a4 pins: forgectrl 522cdb2 (0.1.1), grblhal-glowforge b951f3a (0.1.1), forgefirm-app 1511336 (0.1.22+git)
The audit remediation's commits and the bench session's fixes in the
three repositories, pushed and bench-proven on image 20260902144848
(built from the same commits through a local pin overlay). Each PV
moves with its SRCREV; the manifest records the PV.
2026-09-02 19:00:12 -04:00
ScottW514 82c3596d6a BRINGUP: item 11, the B-16 residual and its actual plan 2026-09-02 18:59:07 -04:00
ScottW514 867136d77e BRINGUP: the Next work list runs 1 to 11, and the closing paragraph closes it 2026-09-02 18:55:25 -04:00
ScottW514 b99c92e047 BRINGUP: the audit's deferred findings as Next work item 11, with the plan for each 2026-09-02 18:51:03 -04:00
ScottW514 5cb18e64ab BRINGUP: the fielded bootloader and the watchdog path, as measured 2026-09-02 18:46:11 -04:00
ScottW514 38cce236a0 forgetest: image.health proves the boot-armed watchdog is active at 60 s
The hardware watchdog is armed by the bootloader and kept fed by the
kernel core; nothing in userspace opens it. Until now nothing in the
catalog read that it is running. image.health reads watchdog0's state,
timeout and bootstatus (CONFIG_WATCHDOG_SYSFS, meta-openglow) and
requires active at 60 s; bootstatus goes into the evidence, so a campaign
that follows a watchdog reset says so.

Bench: on 2026-09-02 the forced hang ended in the factory recovery after
the 60 s timeout; the register read WCR 0x771f. The sysfs view itself
arrives with the next image, so this check is proven there.
Catalog consequence: image.health is always-run.
2026-09-02 18:39:17 -04:00
ScottW514 380f67d666 docs: the forced kernel hang and the two measurements closed
CAMPAIGN-LOG: the forced hang on image 20260902144848 with its times,
the watchdog register read back, U-Boot's recovery branch (the purple
button), the recovery image on the lease, and the power cycle back; the
pad measurement dropped by decision. BRINGUP item 9: both measurements
closed; the pushes, the pins, one build and the campaign on that image
remain.
2026-09-02 18:30:38 -04:00
ScottW514 61f44011aa docs: the pooled bench session's second pass
CAMPAIGN-LOG: the attended set green on image 20260902144848, the three
defects the live runs and the takeover restarts found (the dwell-gap
latch rule, the log drops, the busy-start airflow) and their fixes, and
what is owed. BRINGUP item 9: both queues green; the measurements, the
pushes, the pins, one build, and the campaign on that image are what is
left.
2026-09-02 18:14:18 -04:00
ScottW514 9258dea885 forgetest: fans-quiet also proves a daemon restart on a busy machine returns to idle airflow
A forgectrl started while the kernel is not idle takes the cooldown
airflow (forgectrl's busy-start rule), and on the bench it kept it: after
kernel.fire-line's takeover restarted the daemon with the kernel in the
drill's safe state, the exhaust ran at 6200 rpm on an idle machine until
the daemon was restarted by hand. cooling.fans-quiet-after-motion gains
the case: forgectrl stopped, cnc/disable written, forgectrl started, and
within 90 s the controller must be running with the idle duties applied.
The host replay stubs the init script and holds both outcomes: the idle
duties after the start, and a daemon that keeps the cooldown duties.

Bench: with forgectrl 522cdb2, the busy start logged, idle airflow one
tick later, the duties idle 15 s after the start, PASS.

Catalog consequence: the cooling.* implementation hashes move.
2026-09-02 18:14:07 -04:00
ScottW514 970f10a9e2 forgetest: the dwell-gap latch rule judges the hardware's unlocked window, the live runs keep a trail
laser.emission-witness required the hardware button latch clear in every
sample the engine reported armed, and, after a first fix, in every sample
up to the last nonzero emission count. Both windows were drawn from
lagging signals: the engine's armed flag follows the controller's next
report, and the emission counter latches once per second and reads
nonzero about two seconds past the relock. Both reached into the tail
where the job-end relock sets the button latch by design, and the rule
refused three clean runs on image 20260902144848 (all four sides
burned; the trail shows the latch clear from the press to the relock,
emission through the fourth side, HV_ENABLE's dip in the dwell and its
return).

The rule now uses the window the hardware defines: from the first
emission, in every sample whose readback word shows the laser latch
unlocked, the button-latch bit of that same word must be clear. That
spans the kernel-run gap of the dwell and ends at the relock, and no
lagging flag can misplace it. dwell_gap() is a pure function;
tests/test_laser_dwell.py holds the relocked tail, a set inside the gap,
and a trail without emission. The recorded trail of the third run
replays to a pass (47 unlocked samples, none set).

The live runs keep a per-sample trail in the evidence (TRAIL_FIELDS: the
readback word, the switches, the lock flag, the controller's state and
messages), so a run's timeline can be read back without a rerun.

A fourth run then errored on a name the refactor had removed and one
later check still used; py_compile does not catch it and a live drill
never executes on the host, so the CI job now fails on any undefined
name in the harness (pyflakes).

Catalog consequence: the laser implementation hashes move.
2026-09-02 17:42:07 -04:00
ScottW514 3bb16a48c3 logging: the unix datagram queue holds a daemon's burst
The daemons log to /dev/log with non-blocking datagrams and drop what
the socket will not take. The kernel's default queue for a unix datagram
socket is 10; forgectrl's arm-time burst alone was about 20, so the
lines around every job start were lost, and the daemon reported 18 to
46 dropped per job on the bench. The logging init sets
net.unix.max_dgram_qlen to 512 before rsyslog and the daemons start;
nothing on the image applied sysctl files before.

Bench: set at runtime on 20260902144848 for the rest of the session.
Catalog consequence: a layer change; the meta-forgefirm content hash
moves and nothing inherits on the image that ships it.
2026-09-02 17:26:53 -04:00
ScottW514 774ae52b61 docs: the pooled bench session's first pass, and the PIC read pattern
CAMPAIGN-LOG: the unattended set on image 20260902144848, the six
harness and diagnostic defects it found and their fixes, the numbers,
and the board's state at the end of the pass. BRINGUP: item 9 names the
reset in the hang case and the session under way; the facts bank gains
the PIC read-pattern measurement; Next work item 10 is a pacing of PIC
reads in the kernel.
2026-09-02 17:02:33 -04:00
ScottW514 7605a90946 forgetest: the update drill reads the daemon's reply, and a queue start waits for the fixture probe
update.slots-and-signature's apply section required 200 from
POST /update/apply, where the daemon answers 202 with started, like
every job endpoint, so its first bench run on image 20260902144848
ended before the job did; the cleanup then deleted the staged archive
under the running job. The drill requires 202 and started, and looks
for the daemon's refusal ("archive is not signed with the ForgeFIRM
release key"). Bench: the apply started, the job ended with that
refusal, PASS.

A queue started 4 s after a forgetest restart ran 7 tests instead of
10. The bench page's /state poll had a fixture probe in flight (an mDNS
answer), probe_fixture stamped its time at its start, and the queue
start read the stale fixture, none, so the three operator tests the
fixture runs in the unattended queue were routed to nobody. The probe
now runs under a lock and is stamped when it completes: a caller that
arrives during a probe waits for its answer. tests/test_fixture.py
holds the race with a slow scripted probe; it fails on the old code.

Catalog consequence: the update implementation hash moves; the runner
change is dev-only.
2026-09-02 17:02:33 -04:00
ScottW514 1fef9c6f51 forgetest: the air-assist offset calibration runs before the heater tools
A queue runs the catalog in registration order among tests with the
same prerequisites, and cooling.aa-offset-calibrate followed
cooling.flow-verify. A flow-verify trial heats the tube water (a no-flow
trial by 17 C on the bench) and the warm slug circulates past the
coolant sensors for minutes afterward; the calibration's stationary gate
passed 44 s after the trial on image 20260902144848 and the edges read
the wave as disagreement.

The calibration is registered first now, with the reason beside it, and
tests/test_cooling_order.py holds the order.

Catalog consequence: the cooling.* implementation hashes move (the suite
file changed).
2026-09-02 16:21:13 -04:00
ScottW514 d5369630d9 forgetest: motion.deadman recovers the hung controller with a reset, then the unlock
The hang case resumed the controller and sent $X alone. The stream
fault raises Alarm 17 (motor fault), which the core treats as a critical
event: $X is refused with error:79 until a soft reset, and the reset is
what the stream takes as the operator's acknowledgment of the fault
(grblHAL 38b450e: the kernel stopped and re-armed, the stale ring
cleared, the producer armed again). The drill's first bench run on image
20260902144848 therefore ended in Alarm.

The drill now records $X before the reset and requires the error:79
refusal, sends the soft reset, unlocks, requires Idle, requires the ring
back at its idle free count (the stale bytes of the interrupted move are
gone), and then jogs. The final assertion read the state off the report
dict returned by wait_idle as a string; it never ran before because the
TIMEOUT test short-circuited it.

Bench: motion.deadman PASS on 20260902144848 (kill respawn 1.2 s, hang
to underrun 0.21 s, $X -> ALARM:17 error:79, reset + $X -> Idle, ring
33521664 of 33521664, jog Jog -> Idle, restart retook supervision).

Catalog consequence: the motion.* implementation hashes move (the suite
file changed) and the set re-ran and passed. The failed first run had
closed the campaign, so the always-required core ran again in the new
one, as the campaign rules require.
2026-09-02 15:55:08 -04:00
ScottW514 b3efab9c46 forgetest: the latch-unlock gate waits for the safety chain to release
The gate that refuses a latch unlock while the chain may hold HV_ENABLE
up (charge_pump_alive or a pulse engine not idle) ran at the start of
phases B, U and K3 of kernel.fire-line, within a second of the previous
phase's run. A run feeds the charge-pump watchdog every 200 ms and the
one-shot holds ALIVE for 0.45 s after the last feed, so the gate read
alive=1 and refused: the first bench run of the gate (forgefirm
64f552fc; the laser_pgood gate before it was vacuous) failed phase B on
image 20260902144848.

wait_hv_off() polls the chain for up to 3 s before it refuses, logs the
release when it was not immediate and records every wait in the
evidence (hv_release_s). require_hv_off and check_hv_off use it. The
bench scripts that copy the gate (fire_test.py per phase,
gate_a_kernel_drills.py K3 after K2) get the same wait.

Bench: kernel.fire-line PASS on 20260902144848 with the chain released
after 0.41 s at each of the three phase boundaries. Host:
tests/test_kernel_suite.py covers release inside the window, a chain
held past it, and a chain already off.

Catalog consequence: the kernel.* implementation hashes move (the suite
file changed); the kernel set re-ran and passed.
2026-09-02 15:39:24 -04:00
ScottW514 133b61a062 forgetest: image.health matches the kernel by identity, hash aside
The manifest lists the modules directory without the kernel's
LOCALVERSION_AUTO hash (the hash does not reproduce across a re-patch of
the same source, so the image manifest strips it). image.health still
compared the full running release against that list and failed on the
first post-flash run of image 20260902144848 with the kernel
6.12.20-fslc-fslc-g72a0b1431a9d against the manifest's 6.12.20-fslc-fslc.

kernel_ident() strips the same suffix from both sides, so a manifest with
or without the hash matches the running kernel, and a different base
release still fails. tests/test_image.py covers both forms.

Catalog consequence: only image.health's own implementation hash moves;
it is an always-run test, so no inherited result is affected.
2026-09-02 15:27:36 -04:00
ScottW514 75c4a65444 CAMPAIGN-LOG: the audit remediation, host-tested and committed locally 2026-09-02 09:59:40 -04:00
ScottW514 88ec984e28 Audit follow-through: runbook, bench tools, recipes, release tooling
BRINGUP describes the present: the 54-test catalog and its seven-test
always core, the tier counts, the shipped low-temperature gates, the
density floor ($35 = 10), the two local core commits, the ffboot env
write, the aa-offset route, the current bench image, and the bench
measurements the audit asks for (pooled into the next session). The
workstation shell notes and every em dash are gone.

forgetest: the takeover waits for the cloud client too (found by its
command line); the unauthenticated /boot probe names the endpoint's
parameter; the UI prose is American English. Recipes: forgetest
fetches its package directory and init script only and drops
__pycache__ at unpack; the dev image no longer re-adds forgectrl; the
release image's remove list drops the gfui-client the BSP no longer
has; the platform identity strips the kernel's local-version hash
from the modules directory name, so a re-patched kernel keeps its
fingerprints. grblhal restart is stop then start. release.sh --dev
packs the dev image. fixture.sh refuses a readable env file.

Bench tools: the live-fire drills measure the lid-IR baseline before
every run and point at the fire-watch thresholds the engine reads;
one thermistor conversion (gfbench.degc) serves every drill; the six
dated measurement records leave the tool directory; feeder.c names the
two sysfs writes its caller makes.

Host tests: forgetest 258 pass; the coverage lint reports no uncovered
path across 54 tests. Acceptance: forgectrl.auth covers the /boot
probe; update.* cover ffboot and the manifest identity; the runbook
and bench-tool changes have no catalog consequence.
2026-09-02 09:51:23 -04:00
ScottW514 6002da8d12 ffboot: probe slots without a journal replay
A read-only mount of an ext4 slot still replays its journal, which
writes to the partition the probe was not meant to touch. The probe
mounts with noload: the journal is left as it is and the slot's
content is read as it stands.

Acceptance: update.slots covers ffboot -l (covers map ffboot/**).
2026-09-02 09:12:01 -04:00
ScottW514 0f28427c22 bench and catalog: the controller's cancel messages spell canceled; the needles follow 2026-09-02 08:58:59 -04:00
ScottW514 3b21057dfe BRINGUP: the audit remediation follow-through is a Next work item 2026-09-02 08:22:03 -04:00
ScottW514 b182a5ab0e acceptance: helper imports move the fingerprints, and the update test verifies a foreign signature
A test's fingerprint covered its own text and its module's shared text
only, so a judge imported from a sibling suite module (laser.py takes
its motion judges from motion.py) could change without moving the
fingerprints of the tests that call it. The shared text of every sibling
module a module imports now rides along, transitively; unit test.

update.slots-and-signature claimed to refuse a tampered signature but
fed fwup one garbage file. It now makes a throwaway key pair on the
machine, signs a tiny archive, checks that the archive verifies with its
own key and fails against the shipped release key, and asks the update
job to apply it without confirm_unsigned: the job refuses it for its
signature before touching the slot.
2026-09-02 08:21:01 -04:00
ScottW514 b132965e15 docs: the runbook's settings inventory and the bench page's help links are current
BRINGUP named nine cool_* tunables (there are thirty), a cool_fire_ir_delta
key nothing reads, and panel source files that no longer exist; the
inventory is the present one. The bench page's help popovers pointed at a
documentation host and paths that do not exist; they open the site.
2026-09-02 08:15:12 -04:00
ScottW514 d9a0990510 release: the first release is 0.0.1, and a beta cannot be numbered 0.1.0
FORGEFIRM_RELEASE said 0.1.0, the first non-beta number by the settled
rule, so the first cut was either refused as 0.0.1 or shipped as a
non-beta. The recipe now says 0.0.1, and release.sh refuses a version at
or above 0.1.0 while the README carries the beta banner.
2026-09-02 08:08:20 -04:00