bench: the stale-origin drill goes, drills report their verdict, the pgood gate follows the pin

The faultpos live-fire drill armed and commanded a cut at an origin it
called stale to test a refusal the design decided not to gate: its only
outcome was an emission at an unknown position. Removed from the script
and the bench page.

live_fire_drills.py discarded every drill's return value, so the bench
page recorded a failed live-fire drill as OK. The exit status is the
drill's.

laser_pgood is the supply's power-good, high on every healthy machine;
fire_test.py and the K3 drill aborted on it and pgood_probe.py inverted
it. The latch-unlock drills now gate on the safety chain holding HV off
(charge-pump watchdog dead, pulse engine idle), as the kernel suite
does, and the probe reports the pin as the kernel publishes it.

motion.deadman: the controller resumed from its hang recovers on $X and
moves again without a restart (the stream's fault acknowledgment).
This commit is contained in:
ScottW514
2026-09-02 08:08:19 -04:00
parent 8cf2212e63
commit a2bc4233d5
6 changed files with 59 additions and 46 deletions
+3 -3
View File
@@ -207,10 +207,10 @@ TOOLS = [
# -- laser (live) --------------------------------------------------------------
{"id": "live-fire", "title": "LIVE laser drills", "script": "live_fire_drills.py",
"safety": "live", "where": "board", "ported": True,
"args": [_arg("drill", "choice", "witness", "witness / hold / faultpos / ircut / expstop / ctrlstart",
["witness", "hold", "faultpos", "ircut", "expstop", "ctrlstart"]),
"args": [_arg("drill", "choice", "witness", "witness / hold / ircut / expstop / ctrlstart",
["witness", "hold", "ircut", "expstop", "ctrlstart"]),
_arg("power", "int", 1000, "ircut: S value"), _arg("feed", "int", 300, "ircut: F value")],
"desc": "Emission witness, disarm grace in Hold, stale-origin refusal, lid-IR characterization cut, armed "
"desc": "Emission witness, disarm grace in Hold, lid-IR characterization cut, armed "
"kill on the expected-stop path (+ the separate controller restart). The operator's arm press is "
"required for every drill; eye protection, fire watch, extinguisher, exhaust."},
{"id": "resume-dark-lead", "title": "Pause / resume chain timing (dark lead)", "script": "resume_dark_lead.py",
+21 -5
View File
@@ -501,7 +501,8 @@ def _return_x(ctx, delta_mm):
description="SIGKILL of the controller mid-move: the supervisor reaps it, safes (cnc/stop, "
"latch relocked - it never unlocked), and respawns within seconds. SIGSTOP (a "
"hang) mid-move: the ring drains into a kernel underrun (fast halt, latch "
"locked); the hung process is killed and the supervisor respawns. forgectrl "
"locked); the process resumed from the hang recovers on $X and moves again "
"without a restart. forgectrl "
"restart mid-move: the busy controller finishes the move unmanaged and the new "
"daemon retakes supervision at idle. After each drill the head is jogged back "
"by the kernel-measured distance.")
@@ -584,12 +585,27 @@ def deadman(ctx):
"underruns": hw.sysfs_int("cnc/underruns", 0)}
ctx.log("after SIGSTOP: kernel %s in %s s, latch locked %s, underruns %s -> %s",
kstate, halt_s, latch_locked(), underruns0, ev["sigstop"]["underruns"])
_os.kill(pid1, _signal.SIGKILL) # the hung controller cannot recover itself
# The controller comes back from the hang to find its stream
# faulted and the core alarmed. An unlock ($X) is the operator's
# acknowledgment: it must restore a controller that moves again,
# without a restart (the position is not trusted until a re-home,
# so the move is a jog).
_os.kill(pid1, _signal.SIGCONT)
ctx.sleep(1.0)
st = g.status_report()["state"]
ev["sigstop"]["state_after_cont"] = st
ctx.log("controller resumed: state %s", st)
g.command("$X")
g.command("$J=G91X-5F1200")
peak, states, st = wait_idle(ctx, g, 15)
ev["sigstop"]["recovery_states"] = states
ctx.check("TIMEOUT" not in states and not st.startswith("Alarm"),
"the controller did not move again after $X (states %s)", states)
ctx.check(kstate == "underrun", "the ring did not drain into a kernel underrun (state %s)", kstate)
ctx.check(latch_locked(), "latch unlocked after the underrun")
m2 = wait_running(30, not_pid=pid1)
ev["sigstop"]["respawn"] = m2
ctx.check(m2 and m2.get("pid") != pid1, "supervisor did not respawn after the hang")
m2 = wait_running(10)
ev["sigstop"]["after"] = m2
ctx.check(m2 and m2.get("pid") == pid1, "the recovered controller was replaced (%s)", m2)
ctx.sleep(3)
x1 = _kernel_x_mm(ctx)
_return_x(ctx, (x1 - x0) if (x0 is not None and x1 is not None) else None)