BRINGUP: lid / button / interlock parity is done; items 4 and 12 close with it

Item 16 becomes the record of what the machine does rather than a list of
what is left: both modes cancel a job on a lid or interlock open - running
or paused - and return the head to where the job started with the lid still
open; the park ignores the lid; the arm wait cancels with the reason named;
the button pauses and resumes, with the factory's backtrack and lead in cloud
mode and feed hold / cycle start in GRBL; lid_policy = hold keeps the stock
door behavior for senders that want it. Item 4's mid-job Door hold is that
policy now, and item 12's LightBurn door handling is closed by the cancel -
LightBurn never lives in Door on the default path.

The GRBL resume dwell that was left open is decided against, on measurement
rather than on the mark: the chain re-arms within ~3 ms of the resume while
motion restarts ~219 ms later, so there is nothing to cover.

The two planning files at the tree root are merged and removed, as the audit
plans were. What was durable in them lives here now: the factory's own
reaction timings on 2.6.0-2228, the pause/resume behavior of the safety chain
at pad resolution, and why the hardware button latch is what makes the armed
window honest. The factory session log they were written from is archived
under _RESOURCES/.

Item 15 gains the catalog's current shape: 35 tests after the parity work and
the sweep that merged the tests sharing a setup, with the auto tests left
separate.
This commit is contained in:
ScottW514
2026-08-17 10:02:25 -04:00
parent 674cd1903c
commit 74adaf40b8
+139 -130
View File
@@ -1,6 +1,13 @@
# ForgeFIRM bring-up status & cold-start runbook # ForgeFIRM bring-up status & cold-start runbook
Last updated: **2026-08-15** — **unified logging landed in every repo Last updated: **2026-08-17** — **the machine now reacts to the lid, the
remote-interlock loop and the button the way the factory firmware does, in
both controller modes: a lid or interlock open cancels the job and sends the
head back to where the job started with the lid still open, the button pauses
and resumes, and the software armed window and the hardware button latch agree
by construction. Bench-validated through the acceptance catalog on dev image
`20260817124714` ("Next work" item 16, which closes items 4 and 12).** Before
that: **unified logging landed in every repo
(code-complete, host-verified end to end, pushed, pins bumped): rsyslog (code-complete, host-verified end to end, pushed, pins bumped): rsyslog
is the system logger and the only log writer, every ForgeFIRM process is the system logger and the only log writer, every ForgeFIRM process
emits through syslog under its own program name, each logger has its own emits through syslog under its own program name, each logger has its own
@@ -1748,7 +1755,16 @@ dev image (the confirmation campaign's image).**
451.8 / 455.6 ms; feed period 199.98 ms; the pad-level jog 451.8 / 455.6 ms; feed period 199.98 ms; the pad-level jog
characterization dates from 2026-08-07, sampled at 20 ms through X and characterization dates from 2026-08-07, sampled at 20 ms through X and
Z jogs, ~70/75 samples). It gates nothing anywhere — it is telemetry (`/status` Z jogs, ~70/75 samples). It gates nothing anywhere — it is telemetry (`/status`
`switches.hv_enable`, control-panel "HV enable"). Naming note: the `switches.hv_enable`, control-panel "HV enable").
**Across a pause and a resume** (measured 2026-08-17 at the pads with
`scripts/bench/resume_dark_lead.py`, ~2 kHz through /dev/mem, motion dated
from the kernel counters): a pause stops motion 317 ms after the command and
HV_ENABLE drops with the watchdog 550 ms after it — the pump feed ends with
the run, then t_w runs out — so a pause shorter than about half a second
never drops HV at all. On the resume HV_ENABLE and the watchdog are **back
within ~3 ms** while motion only restarts at ~219 ms: the chain re-arms
roughly 216 ms **before** the first step, so a resumed cut loses nothing to
it and no dark dwell is warranted. Naming note: the
factory design labels this net **E-STOP**, and dated entries below factory design labels this net **E-STOP**, and dated entries below
written before the rename (through the earlier 2026-08-15 records) written before the rename (through the earlier 2026-08-15 records)
call it `estop`/`SW_ESTOP` with the pre-rename polarity (the device tree then declared the pin active-high, call it `estop`/`SW_ESTOP` with the pre-rename polarity (the device tree then declared the pin active-high,
@@ -1760,6 +1776,25 @@ dev image (the confirmation campaign's image).**
line was misread as an e-stop input, and a real e-stop belongs in the line was misread as an e-stop input, and a real e-stop belongs in the
lid-switch chain (`docs/SAFETY.md`). Doors/door1/door2 stay stable lid-switch chain (`docs/SAFETY.md`). Doors/door1/door2 stay stable
during motion. during motion.
- **Factory job behavior on the lid and the button, measured on 2.6.0-2228**
(bench session 2026-08-16, log archived at
`_RESOURCES/factory-session-20260816/`; this is what ForgeFIRM's parity
policy reproduces, "Next work" item 16). Lid open mid-print: `cnc/stop`
5-6 ms after the edge, decel to idle in 86-91 ms, the return-home park
starting ~300-340 ms after the edge and running to completion **with the lid
still open**, the job reported `:cancelled`. A cancel from the app takes the
same path. The button pauses a print — controlled stop, then a 2000-tick
laser-off backtrack — and resumes it with a 1950-tick laser-off lead; the
button **flashes white while paused** (operator-observed). A lid open while
paused cancels the job and parks from where it stands. The lens hunt is not
lid-gated.
- **The hardware button latch is what makes the armed window honest.** A lid
open SETs it (set-dominant), and it stays SET until the lid is closed, the
SoC lock is released **and the button is pressed** (`docs/SAFETY.md`). So a
policy that cancels the job on a lid open and re-arms only through a fresh
button press keeps software and hardware in agreement by construction; one
that resumes a job after a lid open leaves the beam blocked in hardware while
software believes it is armed.
- Machine identity from OCOTP nvmem: HW_OCOTP_MAC0 is the serial, - Machine identity from OCOTP nvmem: HW_OCOTP_MAC0 is the serial,
base-23-encoded to the factory hostname — fuse-verified on the bench base-23-encoded to the factory hostname — fuse-verified on the bench
against the factory label. The bench machine's actual values are against the factory label. The bench machine's actual values are
@@ -2392,11 +2427,10 @@ dev image (the confirmation campaign's image).**
approaches are near-silent** — belt compliance turns slow-speed approaches are near-silent** — belt compliance turns slow-speed
skipping into sub-threshold grinding — so any contact-sensing skipping into sub-threshold grinding — so any contact-sensing
scheme must strike fast. scheme must strike fast.
4. **Controller safety mapping — IMPLEMENTED 2026-08-13; the mid-job 4. **Controller safety mapping — DONE.** The mid-job Door hold described
Door hold described here is superseded by the factory-parity policy of here is the `lid_policy = hold` path; the default is the factory-parity
item 16 (lid = cancel + return to the job start; Door hold only with cancel of item 16 (lid or interlock = cancel + return to the job start),
`lid_policy = hold`) — bench validation bench-validated 2026-08-17 (`grblHAL-glowforge/src/glowforge_switches.c`). The
pending** (`grblHAL-glowforge/src/glowforge_switches.c`). The
controller reads EV_SW with `EVIOCGSW` from the protocol thread's controller reads EV_SW with `EVIOCGSW` from the protocol thread's
realtime hook (no grab — forgectrl polls the same device) and maps: realtime hook (no grab — forgectrl polls the same device) and maps:
- **doors (bit 3) not closed, or interlock (bit 5) loop open → - **doors (bit 3) not closed, or interlock (bit 5) loop open →
@@ -2801,20 +2835,14 @@ dev image (the confirmation campaign's image).**
with the drop both times, i.e. HV_ENABLE = DOORS_OK · WDOG_ALIVE with the drop both times, i.e. HV_ENABLE = DOORS_OK · WDOG_ALIVE
observed live. Full write-up of the chain: `docs/SAFETY.md` observed live. Full write-up of the chain: `docs/SAFETY.md`
(+ `docs/img/safety-chain.svg`). (+ `docs/img/safety-chain.svg`).
12. **LightBurn door-open handling — further issues (found 2026-08-15, 12. **LightBurn door-open handling — CLOSED by item 16.** The lid no
details pending).** With image 20260815154622 (grblHAL a9446fe: door longer parks a job in `Door` on the default policy: it cancels the job,
signal hidden while idle/jog/homing) LightBurn connects again after ends the sender's stream with a clean reset and returns the head to the
an idle lid cycle, but the same bench session turned up other job start, so LightBurn never lives in `Door` and the Resume convention
problems around lid opening in LightBurn that were not characterized it used to need is gone. The `Door` residency that remains under
on the spot. To be detailed and reproduced in a dedicated testing `lid_policy = hold` is covered by `motion.lid-policy-hold` (lid parks the
session: symptoms, whether they involve the mid-job Door hold / job, cycle start after the lid closes finishes the move with its position
Resume path, Start-with-lid-open, or the sender's own handling of the intact), bench-validated 2026-08-17.
`Door` state, and what the controller reports at each step. Until
then the door change stands as partially validated (item 4).
**2026-08-16:** the default mid-job lid path is now the factory cancel
(item 16), so LightBurn no longer lives in `Door` at all; retest the
symptoms with the item-16 tests, and only chase what remains under
`lid_policy = hold`.
13. **uSDHC pad strength brought to the factory values (DTS change 13. **uSDHC pad strength brought to the factory values (DTS change
2026-08-15, bench validation pending — ships with the next full image 2026-08-15, bench validation pending — ships with the next full image
flash, per the batched kernel/BSP rule).** Trigger: one flash, per the batched kernel/BSP rule).** Trigger: one
@@ -3001,112 +3029,93 @@ dev image (the confirmation campaign's image).**
(record in "Release acceptance" above). The flash of `20260816191951` (record in "Release acceptance" above). The flash of `20260816191951`
exposed the layer-hash over-invalidation (a component pin bump counted exposed the layer-hash over-invalidation (a component pin bump counted
as a platform change; fixed - pins in `<recipe>-pin.inc`, left out of as a platform change; fixed - pins in `<recipe>-pin.inc`, left out of
the layer content; record in "Release acceptance" above). Remaining: the layer content; record in "Release acceptance" above). The catalog
(a) the confirmation campaign - a **full** one, on the **first image has since grown to **35 tests** (item 16's parity work, then a sweep
built with the pin files** (that build is a platform change against that merged the tests sharing a setup: `kernel.fire-line` runs A/B/U
every result so far; after it, a component pin bump re-requires only and the mid-ramp unlock behind one takeover, `laser.armed-kill` covers
the tests covering that component) - the tool on the bench is the the expected stop and a SIGKILL on one scrap setup,
tree, but a hot-patched image is not the image that ships; (b) the `laser.pause-resume-lid-cancel` pauses, resumes and then cancels one
bench-tab ports are **code-complete 2026-08-16** (every board-runnable armed burn, and `cloud.lid-interlock-abort` runs the lid and the
tool is ported: the scope tools, the flow characterization family, interlock as two prints; the 17 `auto` tests were left separate, since
the escalation drill, the live drills; record in "Release acceptance" merging them buys no operator time and costs failure isolation).
above) - their bench validation rides the same next dev image; (c) Every board-runnable bench tool is ported to the page, including
the first release runs the full campaign and commits `resume_dark_lead.py`. Remaining: the first release runs the campaign
`releases/v<version>/acceptance.json` - **not yet: no release is and commits `releases/v<version>/acceptance.json` - **not yet: no
cut.** release is cut.**
16. **Lid / button / interlock parity with the factory firmware — CODE-COMPLETE
and host-verified 2026-08-16, bench validation pending.** Both controller 16. **Lid / button / interlock parity with the factory firmware — DONE,
modes now react to the lid, the interlock loop and the button the way the bench-validated 2026-08-17 on dev image `20260817124714`.** Both controller
factory daemon does (its behavior was decoded and then observed on the modes react to the lid, the remote-interlock loop and the button the way the
bench machine booted into factory 2.6.0-2228 the same day: lid open factory daemon does. The factory behavior was decoded and then recorded on
mid-print → `cnc/stop` 5 ms after the edge, immediate return to the job the bench machine booted into factory 2.6.0-2228; that session's log is
start with the lid still open, `:cancelled`; app cancel the same path; archived under `_RESOURCES/factory-session-20260816/` (with a README indexing
button → pause with a 2000-tick laser-off backtrack, resume with a its five prints) and its measured numbers are in the facts bank above.
1950-tick laser-off lead; lid while paused → cancel + park). - **What the machine does, both modes.** Lid or interlock open during a job,
- **GRBL mode** (`grblHAL-glowforge/src/glowforge_switches.c`, running or paused: motion stops within milliseconds of the edge, the job is
`glowforge_laser.c`): the arm wait cancels on lid or interlock (relock, **cancelled and not resumable**, the head returns to the position the job
clean soft reset - no alarm, reason reported; a press with the lid open started from **with the lid still open**, the kernel laser latch relocks and
never arms); the the armed window closes. The next job re-arms with a button press — the same
button is the pause/resume toggle outside the arm wait (feed hold / press the hardware button latch needs, so the software window and the
cycle start; the arming press is consumed and never a pause press); hardware latch agree by construction. The return-home park ignores the lid
lid or interlock mid-job → the core parks the job (planned decel) and and always runs to completion. A lid or interlock open during the pre-run
the driver cancels it — armed window closed, reason reported, soft button wait cancels the job with the reason named. A lid open during a hunt,
reset from the parked state (position kept, no alarm; the sender sees homing, a jog or at idle is ignored. The button pauses and resumes a job:
the banner), then a driver-enqueued `G53 G0` back to the position the in cloud mode with the factory's laser-off backtrack and resume lead
job started from with the door hidden and the latch locked; the (`cloud_pause_backtrack_ticks` 2000 / `cloud_resume_lead_ticks` 1950), in
`lid_policy` setting (`cancel` default / `hold` = stock door hold) GRBL mode as feed hold / cycle start — the kernel refuses a backtrack on a
selects it. Job start = machine position at the Idle → Cycle live-streamed ring, so a resumed GRBL cut picks up where the deceleration
transition. Test hook: `GF_SWITCH_FILE` (file-backed EV_SW word for ended. A pause is not a cancel: the latch stays unlocked and the armed
null-sink builds). window open across it. `lid_policy = hold` selects stock grblHAL door
- **Cloud mode** (`python3-gfhardware/gfhardware/machine.py`, behavior (park in Door, cycle start resumes) instead of the cancel.
`Glowforge-Utilities` basemachine): interlock joins the lid in every - **GRBL** (`grblHAL-glowforge/src/glowforge_switches.c`, `glowforge_laser.c`):
gate; the switch thread wakes the run loop on the edge (stop within the arm wait cancels on lid or interlock with a clean soft reset — no alarm,
milliseconds, level read as backstop); the park ignores the lid and reason reported — and a press with the lid open never arms; the button is
the cancel flag; a hunt ignores the lid; a job refused at start ends the pause/resume toggle outside that wait, the arming press consumed so it
`:cancelled`; the button pauses/resumes a print exactly as the factory is never also a pause; a lid or interlock open mid-job parks the job through
(kernel `resume -2000` / `resume 1950`, `print:paused` / `print:resumed`; the core's door state (planned deceleration, spindle off, position kept) and
`cloud_pause_backtrack_ticks` / `cloud_resume_lead_ticks` settings); the driver then cancels it, resets from the parked state and enqueues a
hunt honors the cancel flag; every job's terminal event is logged. `G53 G0` back to the job start with the door hidden and the latch locked.
- **Proof so far (host):** `laser_arm_test` (17 new checks), The job start is the machine position at the Idle → Cycle transition.
`laser_lifecycle_test.py` (button-wait, lid/interlock in the wait, `GF_SWITCH_FILE` is the file-backed EV_SW word that lets null-sink builds
button toggle, lid/interlock cancel + return to X=0 without alarm, drive these edges in CI.
`lid_policy=hold`), `python3-gfhardware/tests/test_machine_lid_button.py` - **Cloud** (`python3-gfhardware/gfhardware/machine.py`, `Glowforge-Utilities`):
(22 cases), gfutilities tests (58), forgetest unit + coverage lint; the interlock joins the lid in every gate; the switch thread wakes the run
forgectrl builds clean with the three new settings and panel cards. loop on the edge, with the level read kept as a backstop; the park ignores
- **First bench run (dev image 20260817000107, 2026-08-17 00:26 UTC):** the lid and the cancel flag and clears the ring before it moves, so nothing
`laser.lid-cancel-mid-fire` FAILED - the beam stopped and the job was of the abandoned job plays ahead of it; a hunt ignores the lid; a job refused
cancelled as designed, grbl reported "returned to the job start" with at start ends `:cancelled`, never `:completed`; the button pauses and resumes
0.000 mm drift, but the head never moved: the kernel counters stayed at a print exactly as the factory does (`print:paused` / `print:resumed`), and a
+1440 counts (27 mm, where the lid opened), and the baseline's return lid, interlock or service cancel while paused cancels from where it stands.
jog then moved 54 mm and hit the left rail (counters -1442). Root cause - **No resume dwell.** The GRBL resume was suspected of losing its first ~90 ms
in the stream engine, not the cancel policy: the park's `cnc/run` landed to the HV_ENABLE re-arm. Measured on the pads instead
while the kernel was still playing the hold's queued tail (state (`scripts/bench/resume_dark_lead.py`, numbers in the facts bank): the chain
`running`) - the request was refused with EPERM and `ship_pass` treated is back within ~3 ms of the resume and motion only restarts ~219 ms later, so
"refused, kernel running" as started; the kernel then hit its own there is nothing for a dark dwell to cover and none was added.
end-of-data and idled with the park bytes stranded in the ring, and the - **Proof.** Host: `laser_arm_test`, `laser_lifecycle_test.py` (button wait,
NEXT run (the baseline jog) played them first (stale 27 mm -X) plus lid and interlock in the wait, button toggle, cancel + return without alarm,
the jog. Fixed (grblHAL-glowforge): a refused run on a busy kernel stays `lid_policy=hold`), `python3-gfhardware/tests/test_machine_lid_button.py`,
*pending* and is re-issued the moment the kernel reads idle the gfutilities suite, and the forgetest unit tests + coverage lint. Bench,
(`pending_pass`); a soft reset no longer `stop`s a kernel that is only through the acceptance catalog: `motion.button-hold-resume`,
draining a completed stream, and after a mid-motion reset the unplayed `motion.lid-cancel-home` (cancel from Run and from a hold),
residue is cleared (`lseek 1`) once the stop has played out, before `motion.interlock-cancel-home`, `motion.lid-policy-hold`,
any new bytes ship or the device changes hands; the cancel path waits `cloud.lid-interlock-abort`, `cloud.lid-during-button-wait`,
for the kernel drain before the reset. forgetest: the two lid-cancel `cloud.hunt-lid-open`, `cloud.pause-resume`, `cloud.pause-cancel-paths`,
tests now check the KERNEL counters returned (grbl's belief is not `cloud.gfhome-homing` and `cloud.mode-switch` all PASS 2026-08-17; the live
proof), and the baseline reports unplayed ring bytes as a leftover and arm-wait, mid-burn lid cancel, expected stop and armed-kill drills passed
refuses to jog while any exist. To re-run: `motion.lid-cancel-home` the same day (`laser.arm-wait-lid`, `laser.emission-witness`,
first, then `laser.lid-cancel-mid-fire`. `laser.disarm-in-hold`, and the mid-burn lid cancel that the stream-engine
- **Bench validation pending (acceptance catalog):** `laser.arm-wait-lid`, fix below made honest).
`motion.button-hold-resume`, `motion.lid-cancel-home`, - **The stream-engine defect this work found and fixed.** A mid-burn lid
`laser.lid-cancel-mid-fire` (live), `cloud.lid-abort` (live), cancel reported a return the machine never made: the park's `cnc/run` landed
`cloud.lid-during-button-wait`, `cloud.hunt-lid-open`, while the kernel was still playing the hold's queued tail, was refused with
`cloud.pause-resume` (live). Items 4 and 12 above are superseded by EPERM, and "refused, kernel running" was taken for a start — the kernel then
this policy (the mid-job Door hold is no longer the default path); idled with the park bytes stranded, and the next run played them first. Fixed
close them with these tests. in `stepper_stream.c`: a refused run stays *pending* and is re-issued the
Bench 2026-08-17 (dev image 20260817014132): `cloud.lid-abort` and moment the kernel reads idle; a soft reset never stops a kernel that is only
`cloud.lid-during-button-wait` PASSED; `cloud.hunt-lid-open` reported draining a completed stream; a mid-motion reset clears the unplayed residue
FAIL for a harness defect - the hunt had completed with the lid open, once the stop has played out, before any new bytes ship or the device changes
but the test then insisted on switching back to GRBL while the hands; the cancel path waits for the drain before the reset. The lesson is in
service was still re-finding the head after the lid closed (`409 the catalog: the lid tests check the **kernel counters**, not grblHAL's
machine is not idle`). Reworked: the cloud job tests now run **in belief about them, and the baseline refuses to jog while unplayed ring bytes
cloud mode and stay there** (`enter_cloud` reuses a live session, exist.
pid-scoped from the client's own websocket lines; the switch is - Items 4 and 12 above are closed by this policy.
made once from GRBL and declared to the baseline; `wait_quiet`
waits the service's follow-up moves out; the hunt test restarts the
cloud client through the supervisor's stop/start lever for a fresh
connect; every print is judged by its own `print [id]: finished`
line), the baseline is mode-aware (cloud mode owns its kernel
config, lamp, and counters; `controller_mode` is never restored as
a bare setting - that had desynced the persisted mode from the live
one), and the page has an **Ignore prerequisites** switch so any
test can be started alone. Proof: `tests/test_cloud_suite.py` (15
cases, the four tests replayed on the bench's own gfcloud excerpts
+ the run loop's pause/resume lines), baseline/server tests, and a
bench drill of `enter_cloud` (reuse) and of the fresh connect +
hunt detection + quiet wait against the live machine (hunt
`:completed`, 3 follow-up motions, quiet at 41 s). Left for the
operator: `cloud.hunt-lid-open` with the lid actually open,
`cloud.pause-resume` (live print + two presses). Still to observe once on the bench: the
~90 ms HV_ENABLE re-arm gap on a GRBL resume (whether a dark dwell
lead is wanted), the app's rendering of `print:paused`, and a lid open
during the return-to-start motion (should be ignored).