mirror of
https://github.com/openglow-org/forgefirm.git
synced 2026-09-28 09:11:11 -07:00
forgetest: image.health asserts the SDMA clock enable count; the gated engine recorded
The check reads /sys/kernel/debug/clk/sdma/clk_enable_count and fails below 1, naming the fault directly instead of the free-space mismatch it surfaced as after the 150 s settle. BRINGUP item 21 describes the fault and the fix; CAMPAIGN-LOG carries the record of how it was found. Covers: image.health already maps kernel-module-glowforge/** and linux-fslc/**.
This commit is contained in:
+27
-5
@@ -1530,7 +1530,8 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
|
|||||||
`/sys/fs/pstore` mounts with ramoops registered, `/dev/dri/renderD128`
|
`/sys/fs/pstore` mounts with ramoops registered, `/dev/dri/renderD128`
|
||||||
is present and both cameras stream through the GPU demosaic, Wi-Fi
|
is present and both cameras stream through the GPU demosaic, Wi-Fi
|
||||||
associates with `regulatory.db` loaded, the switches sit on `event0`,
|
associates with `regulatory.db` loaded, the switches sit on `event0`,
|
||||||
31 modules load and no DMA channel is held by anyone. Two cosmetic dmesg
|
31 modules load and no DMA channel is held by anyone (which, as the
|
||||||
|
campaign later showed, was the problem: see below). Two cosmetic dmesg
|
||||||
lines came with it: spi-imx reports the absent DMA channel at ERR level
|
lines came with it: spi-imx reports the absent DMA channel at ERR level
|
||||||
and runs PIO, and `consoleblank=0` (uEnv) is an unknown parameter without
|
and runs PIO, and `consoleblank=0` (uEnv) is an unknown parameter without
|
||||||
a virtual console. The crash record is proven: a forced `sysrq-c`
|
a virtual console. The crash record is proven: a forced `sysrq-c`
|
||||||
@@ -1543,9 +1544,8 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
|
|||||||
|
|
||||||
A second round rides the next image, host-proven and unflashed: the
|
A second round rides the next image, host-proven and unflashed: the
|
||||||
kernel is UP (`SMP` off) with performance as its only cpufreq governor;
|
kernel is UP (`SMP` off) with performance as its only cpufreq governor;
|
||||||
spi-imx no longer logs the absent DMA channel (patch 0014) and the
|
spi-imx no longer logs the absent DMA channel (patch 0014);
|
||||||
module no longer logs a run request on an empty ring (grblHAL treats it
|
`consoleblank=0` is gone from the boot
|
||||||
as the ordinary race it is); `consoleblank=0` is gone from the boot
|
|
||||||
arguments; only `wl18xx-fw-4.bin` and `wl18xx-conf.bin` ship for the
|
arguments; only `wl18xx-fw-4.bin` and `wl18xx-conf.bin` ship for the
|
||||||
WL1805 (the current factory image's set); IPv6 is on end to end
|
WL1805 (the current factory image's set); IPv6 is on end to end
|
||||||
(distro feature, `udhcpc6` from the `wlan0 inet6` stanza, forgectrl,
|
(distro feature, `udhcpc6` from the `wlan0 inet6` stanza, forgectrl,
|
||||||
@@ -1565,7 +1565,29 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
|
|||||||
the first Advertise it sees. With that access point's RA and DHCPv6
|
the first Advertise it sees. With that access point's RA and DHCPv6
|
||||||
disabled (as on the other two) the board took the firewall's lease, and
|
disabled (as on the other two) the board took the firewall's lease, and
|
||||||
every service answered on the global address from another VLAN: IPv6 is
|
every service answered on the global address from another VLAN: IPv6 is
|
||||||
on end to end. Left: the item-16 drill and the campaign.
|
on end to end.
|
||||||
|
|
||||||
|
The campaign on dev 20260824201945 then found what every check above had
|
||||||
|
missed: the machine cannot move on any image since the trim. The SDMA
|
||||||
|
engine's `ipg`/`ahb` clocks are enabled only while a dmaengine client
|
||||||
|
holds a channel; imx-sdma leaves them off after probe, and glowforge.ko
|
||||||
|
takes its channel through the SDMA API patch without touching them. The
|
||||||
|
ecspi2 `dmas` had been the only clock holder since the first image, by
|
||||||
|
accident. With the block gated every channel-0 transfer completes at
|
||||||
|
once and moves nothing, so the ring reads back the bounce page: the
|
||||||
|
supervisor's probe logs `cannot start the probe run` at every spawn
|
||||||
|
(`cnc/run` returns -ENODATA because the head sync reads the tail it just
|
||||||
|
published), `cnc/free` exceeds the ring, and `/status` reports a position
|
||||||
|
that never moved. `image.health` failed on the free check after the
|
||||||
|
150 s settle timeout, which is how it surfaced (CAMPAIGN-LOG 2026-08-24,
|
||||||
|
"the SDMA clocks, held by nobody"). The fix, host-proven and unbuilt:
|
||||||
|
`sdma_get_channel()` enables the clocks and `sdma_put_channel()`
|
||||||
|
releases them (patch 0003, the API header), the module calls put on
|
||||||
|
remove and on the probe unwind, the empty-ring run request logs at ERR
|
||||||
|
level again (it was the only kernel-log trace of the fault), and
|
||||||
|
`image.health` asserts the SDMA clock enable count directly. The ecspi2
|
||||||
|
`dmas` stay deleted. Left: build, flash, the item-16 drill and the
|
||||||
|
campaign.
|
||||||
|
|
||||||
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the
|
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the
|
||||||
stale origin unless homing is required (GRBL mode permits unhomed cutting; the
|
stale origin unless homing is required (GRBL mode permits unhomed cutting; the
|
||||||
|
|||||||
@@ -3879,6 +3879,61 @@ and forgetest all answered on it from a host on a different VLAN, and
|
|||||||
the board reached the IPv6 WAN gateway. IPv6 is on end to end; the
|
the board reached the IPv6 WAN gateway. IPv6 is on end to end; the
|
||||||
stale ULA ages out with its own lifetime.
|
stale ULA ages out with its own lifetime.
|
||||||
|
|
||||||
|
## 2026-08-24: the SDMA clocks, held by nobody
|
||||||
|
|
||||||
|
The first campaign on dev 20260824201945 (`c-20260824204310-6b6d`) stalled
|
||||||
|
on `image.health`: the pre-baseline waited its full 150 s for
|
||||||
|
`motion=verified`, which forgectrl never reported, and the test then
|
||||||
|
failed on `cnc/free 35618816 exceeds the ring less its 32 KiB gap`. The
|
||||||
|
forgectrl log had the shape of it: `liveness probe: ERROR - cannot start
|
||||||
|
the probe run` at every controller spawn on every boot since the first
|
||||||
|
kernel-trim image (164619), and `MOTION OK` with a healthy p2p on every
|
||||||
|
boot before it, the last at 17:47Z on 161618. The kernel log had nothing,
|
||||||
|
because the one line that would have said so had been demoted to
|
||||||
|
`dev_dbg` the same afternoon on the belief that it was grblHAL's benign
|
||||||
|
race.
|
||||||
|
|
||||||
|
The ring's own readbacks named the fault. `cnc/position` read X =
|
||||||
|
0x200000 steps on a machine that had not moved (forgectrl showed 39321.60
|
||||||
|
mm), the head index sat 2 MiB ahead of the tail on an idle ring, and
|
||||||
|
scratch6/7 read back the script's constants (the 0x01ffffff index mask
|
||||||
|
and the PWM sample-register address), which is the start of the channel
|
||||||
|
context, not its scratch registers. Every context fetch was returning the
|
||||||
|
bounce page as last written, not SDMA memory. `/sys/kernel/debug/clk/sdma`
|
||||||
|
confirmed it: `clk_enable_count 0`, prepare 1, the engine unclocked.
|
||||||
|
|
||||||
|
The mechanism: imx-sdma enables the engine's `ipg` and `ahb` clocks only
|
||||||
|
in `sdma_alloc_chan_resources`, for a dmaengine client, and disables them
|
||||||
|
at the end of probe. glowforge.ko takes channel 26 through the SDMA API
|
||||||
|
patch's `sdma_get_channel()`, which returned `&sdma->channel[ch]` and
|
||||||
|
nothing more. Until the trim, spi-imx on ecspi2 held two SDMA channels
|
||||||
|
and so held the clocks; the pulse engine had run on that accident since
|
||||||
|
its first image. The round-1 device tree deleted ecspi2's `dmas` on
|
||||||
|
purpose (the ring as the only SDMA client), and took the last clock holder
|
||||||
|
with it. With the block gated a channel-0 transfer completes at once and
|
||||||
|
moves nothing: the script load, the context load, the head sync and the
|
||||||
|
position fetch all "succeed"; `cnc/run` sees head == tail right after the
|
||||||
|
tail publish and returns -ENODATA; grblHAL treats that as its ordinary
|
||||||
|
race and carries on idle; `verify_sdma_script` passes by construction,
|
||||||
|
because the write copies the script into the same bounce page the read
|
||||||
|
returns. The "47 earlier occurrences" of `cannot start cut; no data
|
||||||
|
enqueued` on 164619 and the "no DMA channel held by anyone" observation
|
||||||
|
were this fault, read as noise. The SDMA RAM firmware is not involved: it
|
||||||
|
never loaded on any image.
|
||||||
|
|
||||||
|
The fix, and what proves it so far: `sdma_get_channel()` enables both
|
||||||
|
clocks and a new `sdma_put_channel()` releases them (patch 0003 and the
|
||||||
|
API header); the module calls put in remove and in the probe unwind; the
|
||||||
|
empty-ring run request logs at ERR level again; `image.health` asserts
|
||||||
|
`clk_enable_count >= 1` directly, ahead of the 150 s settle it would
|
||||||
|
otherwise wait out. The ecspi2 `dmas` stay deleted. Host-proven: the
|
||||||
|
patch round-trips against the kernel tree with 0008 on top, the kernel
|
||||||
|
object compiles with no new warnings, the module compiles under `-Werror`
|
||||||
|
(its modpost waits on the rebuilt kernel's export), the module's host
|
||||||
|
tests and the 270 forgetest tests pass. Owed: the image, then on the
|
||||||
|
bench `clk_enable_count` reading 1, `MOTION OK` from the probe,
|
||||||
|
`cnc/free` at 33521664 idle, a GRBL job, and the campaign.
|
||||||
|
|
||||||
## Superseded status notes
|
## Superseded status notes
|
||||||
|
|
||||||
### Shared machine services — remaining polish, as listed 2026-08-13
|
### Shared machine services — remaining polish, as listed 2026-08-13
|
||||||
|
|||||||
@@ -71,7 +71,8 @@ def fds_of(pid):
|
|||||||
("grblhal-glowforge", "CMakeLists.txt"), ("kernel-module-glowforge", "**"),
|
("grblhal-glowforge", "CMakeLists.txt"), ("kernel-module-glowforge", "**"),
|
||||||
("linux-fslc", "**")],
|
("linux-fslc", "**")],
|
||||||
description="The image that is running is the image the manifest describes, with the "
|
description="The image that is running is the image the manifest describes, with the "
|
||||||
"kernel options, the module and the pulse ring it maps, the daemon ownership, "
|
"kernel options, the module, the pulse ring it maps and the SDMA clocks it holds, "
|
||||||
|
"the daemon ownership, "
|
||||||
"the init ordering, and the file modes the release depends on.")
|
"the init ordering, and the file modes the release depends on.")
|
||||||
def image_health(ctx):
|
def image_health(ctx):
|
||||||
ev = ctx.evidence
|
ev = ctx.evidence
|
||||||
@@ -129,6 +130,19 @@ def image_health(ctx):
|
|||||||
"cnc/free %s exceeds the ring less its 32 KiB gap (%s)", free, ring)
|
"cnc/free %s exceeds the ring less its 32 KiB gap (%s)", free, ring)
|
||||||
ctx.check(hw.sysfs_read("cnc/interlock_circuit") is not None, "cnc/interlock_circuit unreadable")
|
ctx.check(hw.sysfs_read("cnc/interlock_circuit") is not None, "cnc/interlock_circuit unreadable")
|
||||||
|
|
||||||
|
# The SDMA engine runs only while a channel holder keeps its ipg/ahb
|
||||||
|
# clocks enabled; imx-sdma leaves them off after probe and only a channel
|
||||||
|
# allocation turns them on. glowforge.ko takes its channel outside
|
||||||
|
# dmaengine and holds the clocks itself. With the block gated every
|
||||||
|
# channel-0 transfer is a silent no-op: the probe cannot start, and the
|
||||||
|
# ring reads back a stale bounce page (free above the ring size, a
|
||||||
|
# position counter that never moved).
|
||||||
|
clk = (_read("/sys/kernel/debug/clk/sdma/clk_enable_count", "") or "").strip()
|
||||||
|
ev["sdma_clk_enable_count"] = clk
|
||||||
|
ctx.log("sdma clk_enable_count: %s", clk or "(unreadable)")
|
||||||
|
ctx.check(clk.isdigit() and int(clk) >= 1,
|
||||||
|
"SDMA clk_enable_count %r: the engine's ipg/ahb clocks are not held", clk)
|
||||||
|
|
||||||
# 4. forgectrl holds /dev/glowforge and supervises the controller
|
# 4. forgectrl holds /dev/glowforge and supervises the controller
|
||||||
pids = hw.pidof("forgectrl")
|
pids = hw.pidof("forgectrl")
|
||||||
ev["forgectrl_pids"] = pids
|
ev["forgectrl_pids"] = pids
|
||||||
|
|||||||
Reference in New Issue
Block a user