diff --git a/docs/BRINGUP.md b/docs/BRINGUP.md index 439b172..2f35572 100644 --- a/docs/BRINGUP.md +++ b/docs/BRINGUP.md @@ -1447,13 +1447,64 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`. overflowed, and the serial layer drops bytes on a full ring (`serial.c`, the overflow flag is set and never read), so the job's M5 and M2 were lost, the window stayed open until the sender disconnected, and the - following job ran unarmed. Fix, in this order: arm on - `state.on && !laser_ok` (the consent question does not depend on the - previous spindle state); clear the spindle state in `gflaser_disarm`; - consume the RX overflow flag (report the error and drop the client, so - a sender that ignores flow control cannot leave a half job behind). - Each part gets a regression test in the null-sink harness with the fix, - and the catalog's `covers` widens to `glowforge_laser.c` and `serial.c`. + following job ran unarmed. Owed: the driver fix on an image (the arm + decided by the window alone, an RX overrun dropping the overrunning + line whole and aborting the job), one bench drill of each scenario, + and the catalog's `covers` widened to `glowforge_laser.c` and + `serial.c`. +21. **A sender change while a job runs: discussion.** Today a sender that + disconnects mid-job leaves the motion running to the end of what the + controller holds, with the window closed and fire suppressed (the + consent belonged to the displaced session), so the job finishes dark + and the material is left with an unfinished cut. This is the stock + Grbl and grblHAL expectation for the motion: the controller has no + notion of sender presence, executes what its planner and RX ring hold, + and then waits; the core's stream code (`stream.c`, + `stream_disconnect`) only switches streams, with no hold and no + alarm, and senders treat a lost connection as a failed job (LightBurn + stops its own side and resumes nothing). ForgeFIRM adds only the + disarm on top. The open question is whether the disarm should also + feed-hold the job, so a reconnecting sender can press and resume where + the cut stopped instead of finding the head at the end of a dark pass: + a hold parks the head over hot material with the assist air on the run + profile, and the grace then closes the window in Hold as it does today; + running on leaves a clean stop position but wastes the piece. Decide + with the gapless pause and resume item (17), which owns the resume + mechanics. +22. **The flow check while the tube is lit.** The arm-time heater check + starts at the session open, so with a prompt press the tube is lit + for most of its window, and a lit CW window adds about 1.5 C to the + rise (0.5 C at 45 % density) against a 1.6 C margin; on top of that the + engine takes its baseline from one sample while the coolant ADC + carries a common-mode offset of about 1 C that steps in when the + airflow goes to the run profile, steps out when it returns to idle, + and toggles between two levels in between. Together they put an + ordinary job's check within a few tenths of the limit. Owed, in order: + the engine's reading (means for the baseline and the end, the tube's + share taken off from the `hv_current` integral, one coefficient per + power model) on an image; three `flowload t1` runs to show the + engine's rise back in the dark band; a `cooling.*` catalog case with + an armed CW load; and if that is not enough, the void-on-emission + design with the tube as its own flow tracer. + Open question, the offset's source: the timing points at the airflow + drive (the step lands one sample after the fans go to run duty, before + any HV, and lifts when they go idle, long after the tube is dark), but + high-voltage energy coupling into the lines between the sensors and + the ADC is the other candidate and a shared path could show both; a + scope on the two sensor lines through a session, fans and tube + switched separately, decides. +23. **Laser power-good: what the line means.** `cnc/laser_pgood` and its + sampled count are defined in the UAPI (active low, one sample every + ~3.9 ms), the facts bank records that the sampled count reads 0 through + real cutting, and the cooling engine warns + `laser power-good degraded during the armed window` whenever fewer + than half the samples read low, so the warning fires at every session + open and carries no information. Nobody knows what the line reports + on this PSU: whether it is the supply's own power-good, an HV-present + flag, a polarity we have inverted, or unconnected. Owed: the line on a + scope against `hv_current` through an armed cut, its meaning written + into the facts bank and the UAPI, and then either a warning that means + something or no warning. **Deliberately not gated:** an armed GRBL job after an underrun cuts at the stale origin unless homing is required (GRBL mode permits unhomed cutting; the diff --git a/docs/CAMPAIGN-LOG.md b/docs/CAMPAIGN-LOG.md index 19c0239..fc857c0 100644 --- a/docs/CAMPAIGN-LOG.md +++ b/docs/CAMPAIGN-LOG.md @@ -4226,6 +4226,88 @@ warned by the engine at every session open, a separate item. The check window's own baseline and the tube term are the two candidates the fix chooses between; nothing is built yet. +## 2026-08-29: the arm-skip and RX-overrun fix, host-proven + +The driver fix for BRINGUP "Next work" item 20 is written in +grblHAL-glowforge and proven on the host; it is not yet on an image. + +**The arm.** `spindleSetState` now arms on `state.on && !laser_ok`: the +consent question reads the window alone, never the previous spindle state. +`gflaser_disarm` leaves the spindle-state record alone, since it is the +core's own view (`spindleGetState`, the `A:S` field, planner sync) and the +condition no longer depends on it. `tests/laser_arm_test.c` gained case H: +arm through the press, bump the sender generation, `gflaser_poll` closes +the window with the record still on, the next laser-on must run the button +wait again and arm; and case I: a laser-on inside the open window does not +re-prompt. On the old condition case H fails three checks; on the new one +the whole harness passes. + +**The ring.** In `serial.c`, `rx_byte` on a full ring now drops the +overrunning line whole: what the ring already holds of it is unwritten back +to the last newline, the rest is discarded through the line's own newline, +real-time characters keep passing (they are taken before the ring), and +the overrun is latched. The driver's realtime hook takes it once, logs it, +reports `RX overrun: the sender ignored flow control (Bf:); job aborted` to +the sender and enqueues `^X`, the same stop as the lid cancel: controlled +deceleration, latch relocked, alarm. `tests/serial_test.c` (new, in CMake +and CI) pushes 73 lines of 14 bytes, overruns on the 74th with a `?` in +the middle, and checks: one byte free after 1022, the `?` taken, the +overrun reported once and only once, every line that fit delivered whole, +no fragment left behind, the next line after the overrun whole, and the +same with the overrun landing mid-line; 11 checks pass. + +**The null-sink harness** (`laser_lifecycle_test.py`, over TCP, no +hardware) gained two scenarios and passes 15 of 15: `sender-change-mid-job` +(M4, a 5 s move, the socket closed at 1 s with the spindle on, reconnect, +`M3 S100` must prompt and re-arm; the disarm message itself is written while +no client is connected and is discarded, so the re-arm is the evidence) and +`rx-overrun` (an armed job, then 120 lines written at once: the overrun +report arrives, the state goes to Alarm, the window closes, and after `$X` +a clean job arms again). `laser_stream_test.py` still passes. + +**What a sender change means in Grbl terms**, for BRINGUP item 21: neither +Grbl nor grblHAL knows a sender is present; a lost connection leaves the +controller executing what its planner and RX ring hold, then waiting; the +core's `stream_disconnect` only switches streams; senders treat the loss as +a failed job. ForgeFIRM keeps that for the motion and adds the disarm. + +## 2026-08-29: the flow check's reading, host-proven + +The cooling engine's flow check (forgectrl `cool.c`) now reads its rise +from means and takes the tube's share off before the limit; written and +proven on the host, not yet on an image (BRINGUP item 22). + +**The reading.** The baseline is the mean of the settled window the gate +has just verified (15 samples at 1 Hz), and that history restarts when the +run airflow profile is applied, so the window is taken entirely under the +profile and the ADC offset that comes with it sits on both sides of the +rise; the arm-time check therefore starts about 15 s after the session +opens instead of one second after. The end reading is the mean of the +check's last 5 s. The tube's share is one coefficient per power model +(`cool_laser_heat_cw` 3.06e-5, `cool_laser_heat_density` 2.36e-5 C per +raw-second of `pic/hv_current`, the numbers of the 2026-08-29 burst runs) +times the current integral from 15 s before the window to 15 s before its +end (the heat's lag to the sensor; emission later than that has not +arrived, so it is not counted, the safe side), bounded at 3 C so no setting +can subtract the check away. The verdict lines carry the share and the raw +rise: `coolant flow verified (heater rise 11.8 C, dT 10.7 C; laser 1.5 off +13.2)`. The two keys are validated in `main.c` (0 to 2e-4) and described in +`SERVICES.md`; the diagnostics' own flow-verify and flow-calibrate are +untouched, since they run with the tube dark. + +**The proof.** `tests/cool_flow_test.c` (new, in CMake and CI) includes the +engine source against a fake sysfs tree and a fake clock, with a loop model +that carries the run-profile offset (-1 C stepping in at the session open +and toggling 0.6 C), the heater's rise with flow (12 C plateau) or without +(0.4 C per second) and the tube's heat arriving 15 s after emission. Fifteen +checks pass: a dark check with flow is verified with no share and its +baseline is the window's mean; a dark check without flow is SUSPECT; a lit +CW check with flow is verified with about 1.5 C taken off (raw 13.2, judged +11.8); a lit check without flow is SUSPECT (20.2 raw, 18.8 judged); an +absurd coefficient is bounded at 3 C and no flow is still SUSPECT (17.2); +under the density model the density coefficient applies (1.1 off). The +`flowload` drill's verdict parser accepts the new suffix. + ## Superseded status notes ### Shared machine services — remaining polish, as listed 2026-08-13 diff --git a/scripts/bench/laser_lifecycle_test.py b/scripts/bench/laser_lifecycle_test.py index 0a4363f..59bd957 100644 --- a/scripts/bench/laser_lifecycle_test.py +++ b/scripts/bench/laser_lifecycle_test.py @@ -291,11 +291,85 @@ def test_job_window(): s.close() +def test_sender_change_mid_job(): + """Rule 3, the hard case: the sender changes while the job is still + running with the spindle on, so the core never turns the spindle off + between the two sessions. The next laser-on from the new sender must + still prompt: the arm decision reads the window, not the spindle-state + record (a job whose M5 was lost leaves that record on the same way). + M3 after M4 is a state change the core always pushes through + set_state, planner-synced, so it lands once the first job's move ends.""" + s = Session("sender-change-mid-job", disarm_s=60) + try: + send_line(s.sock, "M4 S100", s.log) + send_line(s.sock, "G1 X5 F60", s.log) # 5 s of motion + if not wait_for(s.log, ARMED, 5, s.sock): + fail("[sender-change-mid-job] first laser-on did not arm") + time.sleep(1.0) + s.sock.close() # mid-move, spindle on + time.sleep(0.5) + # The disarm message is written while no client is connected, and + # output with no client is discarded, so the new session cannot + # see it; the re-arm below is the evidence the window closed. + s.sock = s.connect() + before = s.armed_count() + s.send_raw("M3 S100") # laser-on against a closed window + end = time.time() + 15 + while time.time() < end and s.armed_count() != before + 1: + read_avail(s.sock, s.log, 0.2) + if s.armed_count() != before + 1: + fail("[sender-change-mid-job] a laser-on after a mid-job sender change did not " + "re-arm: the arm read the stale spindle state instead of the window") + send_line(s.sock, "M5", s.log) + wait_idle(s.sock, s.log) + print("PASS [sender-change-mid-job]: a laser-on against a window closed mid-job " + "prompted and re-armed") + finally: + s.close() + + +def test_rx_overrun_aborts(): + """A sender that writes past the RX ring (Bf: is the contract) loses + lines, and a job with lines missing is not the job the sender wrote: + the controller reports the overrun, stops the way a ^X stops it + (alarm, window closed) and never runs on with a hole in it. A clean + job after the reset arms again.""" + s = Session("rx-overrun", disarm_s=60) + try: + send_line(s.sock, "M4 S100", s.log) + send_line(s.sock, "G1 X1 F600", s.log) + if not wait_for(s.log, ARMED, 5, s.sock): + fail("[rx-overrun] first laser-on did not arm") + job = "".join("G1 X%d F60\n" % (2 + (i % 2)) for i in range(120)) # ~1.3 KB at once + s.sock.sendall(job.encode()) + if not wait_for(s.log, "RX overrun", 10, s.sock): + fail("[rx-overrun] the overrun was not reported to the sender") + if not s.wait_state("Alarm", 5): + fail("[rx-overrun] the overrun did not stop the job (state %s)" % s.state()) + if s.disarmed_count() < 1: + read_avail(s.sock, s.log, 1.0) + if s.disarmed_count() < 1: + fail("[rx-overrun] the overrun did not close the armed window") + send_line(s.sock, "$X", s.log) + before = s.armed_count() + send_line(s.sock, "M4 S100", s.log) + send_line(s.sock, "G1 X0 F600", s.log) + end = time.time() + 5 + while time.time() < end and s.armed_count() != before + 1: + read_avail(s.sock, s.log, 0.2) + if s.armed_count() != before + 1: + fail("[rx-overrun] a clean job after the overrun did not arm") + wait_idle(s.sock, s.log) + print("PASS [rx-overrun]: the overrun was reported, the job stopped in alarm with " + "the window closed, and the next clean job armed") + finally: + s.close() + + def test_sender_change(): """Rule 3: a reconnected sender must re-arm. The grace is set far beyond the test horizon so only the sender change can close the - window, and the first job ends in M5 so the next M4 is a fresh - laser-on edge (the arm prompt fires on the off->on edge).""" + window; the first job ends in M5 and the next M4 is a fresh laser-on.""" s = Session("sender-change", disarm_s=60) try: send_line(s.sock, "M4 S100", s.log) @@ -651,6 +725,8 @@ def main(): fail("controller binary not found at %s" % BIN) test_job_window() test_sender_change() + test_sender_change_mid_job() + test_rx_overrun_aborts() test_hold_grace() test_button_wait_arms() test_lid_open_in_wait() diff --git a/scripts/bench/live_fire_drills.py b/scripts/bench/live_fire_drills.py index 3e64153..ced241f 100644 --- a/scripts/bench/live_fire_drills.py +++ b/scripts/bench/live_fire_drills.py @@ -1585,7 +1585,7 @@ FLOWLOAD_PRESS_WAIT_S = 300.0 # the operator's button timeout FLOWLOAD_CHANNELS = PCURVE_CHANNELS + (('htr', 'thermal/heater_pwm'),) # The engine's own verdict lines, as /cool/status "reason" carries them. FLOWLOAD_RISE_RE = re.compile( - r'heater rise ([0-9.]+) C(?: \(limit ([0-9.]+), dT ([0-9.]+)\)|, dT ([0-9.]+) C)') + r'heater rise ([0-9.]+) C(?: \(limit ([0-9.]+), dT ([0-9.]+)|, dT ([0-9.]+) C)') class CoolPoller: