Add the sender-change and RX-overrun harness scenarios and the day's records

laser_lifecycle_test.py gains sender-change-mid-job (a laser-on against a
window closed while the spindle was on must prompt again) and rx-overrun
(a job written past the RX ring is reported, stopped in alarm with the
window closed, and a clean job arms after it). The flowload drill's
verdict parser accepts the engine's laser-share suffix.

BRINGUP: item 20 holds only the owed work; item 21 opens the mid-job
sender-change discussion with the Grbl expectation; item 22 is the flow
check under a lit tube; item 23 is the power-good line's meaning.
CAMPAIGN-LOG records the driver fix and the flow-check reading, both
host-proven.

No catalog consequence: harness scenarios and documentation; no runtime
behavior of the release image changes in this commit.
This commit is contained in:
ScottW514
2026-08-29 14:58:21 -04:00
parent b5de7cda9d
commit 9d9f235219
4 changed files with 219 additions and 10 deletions
+58 -7
View File
@@ -1447,13 +1447,64 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
overflowed, and the serial layer drops bytes on a full ring (`serial.c`, overflowed, and the serial layer drops bytes on a full ring (`serial.c`,
the overflow flag is set and never read), so the job's M5 and M2 were the overflow flag is set and never read), so the job's M5 and M2 were
lost, the window stayed open until the sender disconnected, and the lost, the window stayed open until the sender disconnected, and the
following job ran unarmed. Fix, in this order: arm on following job ran unarmed. Owed: the driver fix on an image (the arm
`state.on && !laser_ok` (the consent question does not depend on the decided by the window alone, an RX overrun dropping the overrunning
previous spindle state); clear the spindle state in `gflaser_disarm`; line whole and aborting the job), one bench drill of each scenario,
consume the RX overflow flag (report the error and drop the client, so and the catalog's `covers` widened to `glowforge_laser.c` and
a sender that ignores flow control cannot leave a half job behind). `serial.c`.
Each part gets a regression test in the null-sink harness with the fix, 21. **A sender change while a job runs: discussion.** Today a sender that
and the catalog's `covers` widens to `glowforge_laser.c` and `serial.c`. disconnects mid-job leaves the motion running to the end of what the
controller holds, with the window closed and fire suppressed (the
consent belonged to the displaced session), so the job finishes dark
and the material is left with an unfinished cut. This is the stock
Grbl and grblHAL expectation for the motion: the controller has no
notion of sender presence, executes what its planner and RX ring hold,
and then waits; the core's stream code (`stream.c`,
`stream_disconnect`) only switches streams, with no hold and no
alarm, and senders treat a lost connection as a failed job (LightBurn
stops its own side and resumes nothing). ForgeFIRM adds only the
disarm on top. The open question is whether the disarm should also
feed-hold the job, so a reconnecting sender can press and resume where
the cut stopped instead of finding the head at the end of a dark pass:
a hold parks the head over hot material with the assist air on the run
profile, and the grace then closes the window in Hold as it does today;
running on leaves a clean stop position but wastes the piece. Decide
with the gapless pause and resume item (17), which owns the resume
mechanics.
22. **The flow check while the tube is lit.** The arm-time heater check
starts at the session open, so with a prompt press the tube is lit
for most of its window, and a lit CW window adds about 1.5 C to the
rise (0.5 C at 45 % density) against a 1.6 C margin; on top of that the
engine takes its baseline from one sample while the coolant ADC
carries a common-mode offset of about 1 C that steps in when the
airflow goes to the run profile, steps out when it returns to idle,
and toggles between two levels in between. Together they put an
ordinary job's check within a few tenths of the limit. Owed, in order:
the engine's reading (means for the baseline and the end, the tube's
share taken off from the `hv_current` integral, one coefficient per
power model) on an image; three `flowload t1` runs to show the
engine's rise back in the dark band; a `cooling.*` catalog case with
an armed CW load; and if that is not enough, the void-on-emission
design with the tube as its own flow tracer.
Open question, the offset's source: the timing points at the airflow
drive (the step lands one sample after the fans go to run duty, before
any HV, and lifts when they go idle, long after the tube is dark), but
high-voltage energy coupling into the lines between the sensors and
the ADC is the other candidate and a shared path could show both; a
scope on the two sensor lines through a session, fans and tube
switched separately, decides.
23. **Laser power-good: what the line means.** `cnc/laser_pgood` and its
sampled count are defined in the UAPI (active low, one sample every
~3.9 ms), the facts bank records that the sampled count reads 0 through
real cutting, and the cooling engine warns
`laser power-good degraded during the armed window` whenever fewer
than half the samples read low, so the warning fires at every session
open and carries no information. Nobody knows what the line reports
on this PSU: whether it is the supply's own power-good, an HV-present
flag, a polarity we have inverted, or unconnected. Owed: the line on a
scope against `hv_current` through an armed cut, its meaning written
into the facts bank and the UAPI, and then either a warning that means
something or no warning.
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the **Deliberately not gated:** an armed GRBL job after an underrun cuts at the
stale origin unless homing is required (GRBL mode permits unhomed cutting; the stale origin unless homing is required (GRBL mode permits unhomed cutting; the
+82
View File
@@ -4226,6 +4226,88 @@ warned by the engine at every session open, a separate item. The check
window's own baseline and the tube term are the two candidates the fix window's own baseline and the tube term are the two candidates the fix
chooses between; nothing is built yet. chooses between; nothing is built yet.
## 2026-08-29: the arm-skip and RX-overrun fix, host-proven
The driver fix for BRINGUP "Next work" item 20 is written in
grblHAL-glowforge and proven on the host; it is not yet on an image.
**The arm.** `spindleSetState` now arms on `state.on && !laser_ok`: the
consent question reads the window alone, never the previous spindle state.
`gflaser_disarm` leaves the spindle-state record alone, since it is the
core's own view (`spindleGetState`, the `A:S` field, planner sync) and the
condition no longer depends on it. `tests/laser_arm_test.c` gained case H:
arm through the press, bump the sender generation, `gflaser_poll` closes
the window with the record still on, the next laser-on must run the button
wait again and arm; and case I: a laser-on inside the open window does not
re-prompt. On the old condition case H fails three checks; on the new one
the whole harness passes.
**The ring.** In `serial.c`, `rx_byte` on a full ring now drops the
overrunning line whole: what the ring already holds of it is unwritten back
to the last newline, the rest is discarded through the line's own newline,
real-time characters keep passing (they are taken before the ring), and
the overrun is latched. The driver's realtime hook takes it once, logs it,
reports `RX overrun: the sender ignored flow control (Bf:); job aborted` to
the sender and enqueues `^X`, the same stop as the lid cancel: controlled
deceleration, latch relocked, alarm. `tests/serial_test.c` (new, in CMake
and CI) pushes 73 lines of 14 bytes, overruns on the 74th with a `?` in
the middle, and checks: one byte free after 1022, the `?` taken, the
overrun reported once and only once, every line that fit delivered whole,
no fragment left behind, the next line after the overrun whole, and the
same with the overrun landing mid-line; 11 checks pass.
**The null-sink harness** (`laser_lifecycle_test.py`, over TCP, no
hardware) gained two scenarios and passes 15 of 15: `sender-change-mid-job`
(M4, a 5 s move, the socket closed at 1 s with the spindle on, reconnect,
`M3 S100` must prompt and re-arm; the disarm message itself is written while
no client is connected and is discarded, so the re-arm is the evidence) and
`rx-overrun` (an armed job, then 120 lines written at once: the overrun
report arrives, the state goes to Alarm, the window closes, and after `$X`
a clean job arms again). `laser_stream_test.py` still passes.
**What a sender change means in Grbl terms**, for BRINGUP item 21: neither
Grbl nor grblHAL knows a sender is present; a lost connection leaves the
controller executing what its planner and RX ring hold, then waiting; the
core's `stream_disconnect` only switches streams; senders treat the loss as
a failed job. ForgeFIRM keeps that for the motion and adds the disarm.
## 2026-08-29: the flow check's reading, host-proven
The cooling engine's flow check (forgectrl `cool.c`) now reads its rise
from means and takes the tube's share off before the limit; written and
proven on the host, not yet on an image (BRINGUP item 22).
**The reading.** The baseline is the mean of the settled window the gate
has just verified (15 samples at 1 Hz), and that history restarts when the
run airflow profile is applied, so the window is taken entirely under the
profile and the ADC offset that comes with it sits on both sides of the
rise; the arm-time check therefore starts about 15 s after the session
opens instead of one second after. The end reading is the mean of the
check's last 5 s. The tube's share is one coefficient per power model
(`cool_laser_heat_cw` 3.06e-5, `cool_laser_heat_density` 2.36e-5 C per
raw-second of `pic/hv_current`, the numbers of the 2026-08-29 burst runs)
times the current integral from 15 s before the window to 15 s before its
end (the heat's lag to the sensor; emission later than that has not
arrived, so it is not counted, the safe side), bounded at 3 C so no setting
can subtract the check away. The verdict lines carry the share and the raw
rise: `coolant flow verified (heater rise 11.8 C, dT 10.7 C; laser 1.5 off
13.2)`. The two keys are validated in `main.c` (0 to 2e-4) and described in
`SERVICES.md`; the diagnostics' own flow-verify and flow-calibrate are
untouched, since they run with the tube dark.
**The proof.** `tests/cool_flow_test.c` (new, in CMake and CI) includes the
engine source against a fake sysfs tree and a fake clock, with a loop model
that carries the run-profile offset (-1 C stepping in at the session open
and toggling 0.6 C), the heater's rise with flow (12 C plateau) or without
(0.4 C per second) and the tube's heat arriving 15 s after emission. Fifteen
checks pass: a dark check with flow is verified with no share and its
baseline is the window's mean; a dark check without flow is SUSPECT; a lit
CW check with flow is verified with about 1.5 C taken off (raw 13.2, judged
11.8); a lit check without flow is SUSPECT (20.2 raw, 18.8 judged); an
absurd coefficient is bounded at 3 C and no flow is still SUSPECT (17.2);
under the density model the density coefficient applies (1.1 off). The
`flowload` drill's verdict parser accepts the new suffix.
## Superseded status notes ## Superseded status notes
### Shared machine services — remaining polish, as listed 2026-08-13 ### Shared machine services — remaining polish, as listed 2026-08-13
+78 -2
View File
@@ -291,11 +291,85 @@ def test_job_window():
s.close() s.close()
def test_sender_change_mid_job():
"""Rule 3, the hard case: the sender changes while the job is still
running with the spindle on, so the core never turns the spindle off
between the two sessions. The next laser-on from the new sender must
still prompt: the arm decision reads the window, not the spindle-state
record (a job whose M5 was lost leaves that record on the same way).
M3 after M4 is a state change the core always pushes through
set_state, planner-synced, so it lands once the first job's move ends."""
s = Session("sender-change-mid-job", disarm_s=60)
try:
send_line(s.sock, "M4 S100", s.log)
send_line(s.sock, "G1 X5 F60", s.log) # 5 s of motion
if not wait_for(s.log, ARMED, 5, s.sock):
fail("[sender-change-mid-job] first laser-on did not arm")
time.sleep(1.0)
s.sock.close() # mid-move, spindle on
time.sleep(0.5)
# The disarm message is written while no client is connected, and
# output with no client is discarded, so the new session cannot
# see it; the re-arm below is the evidence the window closed.
s.sock = s.connect()
before = s.armed_count()
s.send_raw("M3 S100") # laser-on against a closed window
end = time.time() + 15
while time.time() < end and s.armed_count() != before + 1:
read_avail(s.sock, s.log, 0.2)
if s.armed_count() != before + 1:
fail("[sender-change-mid-job] a laser-on after a mid-job sender change did not "
"re-arm: the arm read the stale spindle state instead of the window")
send_line(s.sock, "M5", s.log)
wait_idle(s.sock, s.log)
print("PASS [sender-change-mid-job]: a laser-on against a window closed mid-job "
"prompted and re-armed")
finally:
s.close()
def test_rx_overrun_aborts():
"""A sender that writes past the RX ring (Bf: is the contract) loses
lines, and a job with lines missing is not the job the sender wrote:
the controller reports the overrun, stops the way a ^X stops it
(alarm, window closed) and never runs on with a hole in it. A clean
job after the reset arms again."""
s = Session("rx-overrun", disarm_s=60)
try:
send_line(s.sock, "M4 S100", s.log)
send_line(s.sock, "G1 X1 F600", s.log)
if not wait_for(s.log, ARMED, 5, s.sock):
fail("[rx-overrun] first laser-on did not arm")
job = "".join("G1 X%d F60\n" % (2 + (i % 2)) for i in range(120)) # ~1.3 KB at once
s.sock.sendall(job.encode())
if not wait_for(s.log, "RX overrun", 10, s.sock):
fail("[rx-overrun] the overrun was not reported to the sender")
if not s.wait_state("Alarm", 5):
fail("[rx-overrun] the overrun did not stop the job (state %s)" % s.state())
if s.disarmed_count() < 1:
read_avail(s.sock, s.log, 1.0)
if s.disarmed_count() < 1:
fail("[rx-overrun] the overrun did not close the armed window")
send_line(s.sock, "$X", s.log)
before = s.armed_count()
send_line(s.sock, "M4 S100", s.log)
send_line(s.sock, "G1 X0 F600", s.log)
end = time.time() + 5
while time.time() < end and s.armed_count() != before + 1:
read_avail(s.sock, s.log, 0.2)
if s.armed_count() != before + 1:
fail("[rx-overrun] a clean job after the overrun did not arm")
wait_idle(s.sock, s.log)
print("PASS [rx-overrun]: the overrun was reported, the job stopped in alarm with "
"the window closed, and the next clean job armed")
finally:
s.close()
def test_sender_change(): def test_sender_change():
"""Rule 3: a reconnected sender must re-arm. The grace is set far """Rule 3: a reconnected sender must re-arm. The grace is set far
beyond the test horizon so only the sender change can close the beyond the test horizon so only the sender change can close the
window, and the first job ends in M5 so the next M4 is a fresh window; the first job ends in M5 and the next M4 is a fresh laser-on."""
laser-on edge (the arm prompt fires on the off->on edge)."""
s = Session("sender-change", disarm_s=60) s = Session("sender-change", disarm_s=60)
try: try:
send_line(s.sock, "M4 S100", s.log) send_line(s.sock, "M4 S100", s.log)
@@ -651,6 +725,8 @@ def main():
fail("controller binary not found at %s" % BIN) fail("controller binary not found at %s" % BIN)
test_job_window() test_job_window()
test_sender_change() test_sender_change()
test_sender_change_mid_job()
test_rx_overrun_aborts()
test_hold_grace() test_hold_grace()
test_button_wait_arms() test_button_wait_arms()
test_lid_open_in_wait() test_lid_open_in_wait()
+1 -1
View File
@@ -1585,7 +1585,7 @@ FLOWLOAD_PRESS_WAIT_S = 300.0 # the operator's button timeout
FLOWLOAD_CHANNELS = PCURVE_CHANNELS + (('htr', 'thermal/heater_pwm'),) FLOWLOAD_CHANNELS = PCURVE_CHANNELS + (('htr', 'thermal/heater_pwm'),)
# The engine's own verdict lines, as /cool/status "reason" carries them. # The engine's own verdict lines, as /cool/status "reason" carries them.
FLOWLOAD_RISE_RE = re.compile( FLOWLOAD_RISE_RE = re.compile(
r'heater rise ([0-9.]+) C(?: \(limit ([0-9.]+), dT ([0-9.]+)\)|, dT ([0-9.]+) C)') r'heater rise ([0-9.]+) C(?: \(limit ([0-9.]+), dT ([0-9.]+)|, dT ([0-9.]+) C)')
class CoolPoller: class CoolPoller: