forgectrl pinned at the bench-proven GPU path; the session's record

The pin moves to forgectrl 6614833: the five hardware corrections from
the GPU demosaic's first bench session (surfaceless EGL config, ARGB
render targets, GR88 raw import, the ipu_copy stride-fix crop between
the GPU and the CODA, the chroma mirror width). CAMPAIGN-LOG carries
the dated record of how each was found; BRINGUP item 20 now separates
what that session proved (the full GPU -> IPU -> VPU path serving
correct frames, H.264 as valid fragmented MP4 on hardware, texture
limits, slot fit) from what remains (the 140 ms render, the bottom
row, MSE playback, the CPU measurement, frame skip, coexistence, the
campaign).
This commit is contained in:
ScottW514
2026-08-24 09:11:50 -04:00
parent a47c0c7f6d
commit 904dfb4acb
3 changed files with 66 additions and 24 deletions
+26 -23
View File
@@ -1446,29 +1446,32 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
are established. A pause on contact, on the factory's shape, would be are established. A pause on contact, on the factory's shape, would be
the first use. the first use.
20. **Video pipeline offload — bench validation.** Code-complete and 20. **Video pipeline offload — bench validation.** First hardware session
host-proven (CI: `mp4mux_test`; catalog: `camera.h264-stream`), not yet done (2026-08-24, dev image 20260824122014, drill binaries; fixes in
run on hardware: the GC880 GLES2 demosaic (`src/gpu_debayer.c`, forgectrl 6614833). Proven: Mesa etnaviv fits the release slot
dlopen'd Mesa, capture dmabuf in, encoder dmabuf out), the CODA960 (~5 MiB margin); surfaceless EGL and dmabuf import both directions;
H.264 stream (`/cam/h264`, fragmented MP4, panel MSE player, `GL_MAX_TEXTURE_SIZE` 8192 (no tiling even at 8 MP); the full path
~1.5 Mbit/s vs MJPEG's ~9), and the CSI hardware frame skip behind GPU render → IPU stride-fix crop (`src/ipu_copy.c`, the render
`FORGECTRL_STREAM_FPS`. Falls back to the NEON path wherever a piece engine's 64-byte rows and the CODA's round_up(width,16) stride never
is missing, so the existing stream is not at risk. To settle on the meet, so the IPU crops between them, 14 ms, no CPU touch) → VPU
bench, in order: (1) the image builds with Mesa etnaviv inside the encode, `convert: "gpu"`, image correct to within 2 counts of the
200 MiB slot cap (distro `PACKAGECONFIG:pn-mesa`, image adds scalar demosaic; `/cam/h264` serving valid fragmented MP4 on
`libegl-mesa libgles2-mesa libgbm mesa-megadriver`); (2) surfaceless hardware (avc1.424020, ~480 kbit/s on a static bed). Remaining, in
EGL comes up and dmabuf import works both directions (the open logs order: (a) **GPU render time, 140 ms/frame** (~6 fps ceiling; the
say exactly which probe refused); (3) `FORGECTRL_GPU_CHECK` reports a core runs its full 528 MHz, suspicion is the pre-HALTI
max delta of a couple of counts against the scalar demosaic; linear-texture sampling path - decompose per-pass, consider a tiled
(4) `GL_MAX_TEXTURE_SIZE` covers 1944 rows (5 MP single-tile; the shadow copy or fewer chroma taps); (b) the **bottom output row**
8 MP tiling path can only be exercised on an HD machine or by differs from the CPU demosaic (1296 samples, max delta 134; suspect
forcing tiles); (5) coda H.264 rate control and picture quality at the IC's last-line behavior); (c) browser MSE playback of the panel's
1296x972p15, and the measured CPU with a panel H.264 viewer vs the H.264 view; (d) measured CPU with an H.264 viewer vs the 41 % MJPEG
41 % MJPEG baseline; (6) the stream-during-jog coexistence drill baseline; (e) CSI hardware frame-skip exercise; (f) the
rerun with the GPU path active. Switches to strip a suspect layer: stream-during-jog coexistence drill with the GPU path active;
`FORGECTRL_NO_GPU`, `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, (g) `camera.h264-stream` and the full campaign on an image carrying
plus the existing `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` / the fixes. Switches to strip a suspect layer: `FORGECTRL_NO_GPU`,
`FORGECTRL_NO_CACHED_BUFS`. `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, plus the existing
`FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` /
`FORGECTRL_NO_CACHED_BUFS`; `FORGECTRL_GPU_CHECK` tightens the
stream-stats cadence and logs the render-versus-copy split.
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the **Deliberately not gated:** an armed GRBL job after an underrun cuts at the
stale origin unless homing is required (GRBL mode permits unhomed cutting; the stale origin unless homing is required (GRBL mode permits unhomed cutting; the
+39
View File
@@ -3630,6 +3630,45 @@ on the bed, and the two together by one real print. Still open from that
plan: the bench actuator for the lid, interlock and button, and the plan: the bench actuator for the lid, interlock and button, and the
finer coverage maps. finer coverage maps.
## 2026-08-24: the GPU demosaic's first light, and what the probes caught
Dev image 20260824122014 (the first with Mesa etnaviv; release ext4
204,140,544 bytes, ~5 MiB under the slot cap) flashed by the operator;
the drill ran forgectrl builds from `/tmp` against the running image,
each iteration probed over the stream, `FORGECTRL_GPU_CHECK`, and frame
captures diffed against the CPU demosaic on the host.
Five faults found and fixed in one session, each named by a probe log
line or the compare (forgectrl 6614833):
1. `eglChooseConfig` returned nothing: EGL_SURFACE_TYPE defaults to
WINDOW_BIT and the surfaceless platform has no window configs. Ask
for surface type 0.
2. Every fourth output byte was 255: the render engine writes an
XRGB8888 surface's undefined X byte as opaque. Render ARGB8888.
3. The raw import failed etnaviv's stride check (width padded to 16
texels): 2592 bytes is 1296 GR88 texels exactly, not 656 padded
XRGB ones. Import GR88, one texel per Bayer pair.
4. The GPU cannot write the CODA's buffers at all: 64-byte render rows
versus round_up(width,16) strides never meet at these widths. New
`ipu_copy` module: render into the IPU CSC/scaler's wider source
(stride align(w,128)) and let the IPU crop into each encoder over
dmabuf. 14 ms a copy, no CPU touch.
5. The chroma mirror used a quarter-width plane where the plane is
half-width: the right half of both chroma planes clamped to
column 0. The three-frame diff-by-transform analysis on the host
named both this and fault 2.
End state on the bench: `convert: "gpu"` serving MJPEG, the GPU/CPU
compare clean to 2 counts except the bottom row (1296 samples, max
delta 134, unexplained); `/cam/h264` delivering valid fragmented MP4
from the CODA BIT processor (avc1.424020, ~480 kbit/s on the static
bed). Open, measured: the render costs 140 ms a frame against the IPU's
14 (GPU at its full 528 MHz - the suspicion is pre-HALTI linear-texture
sampling), so the GPU path holds ~6 fps until that is run down. The
`getenv` implicit-declaration fix in debayer.c rode along. Bench left
clean; stock service restored.
## Superseded status notes ## Superseded status notes
### Shared machine services — remaining polish, as listed 2026-08-13 ### Shared machine services — remaining polish, as listed 2026-08-13
@@ -2,5 +2,5 @@
# only SRCREV and PV here - the image manifest leaves *-pin.inc out of the # only SRCREV and PV here - the image manifest leaves *-pin.inc out of the
# layer content hash because the component entry already identifies the # layer content hash because the component entry already identifies the
# pinned source (forgefirm-image-manifest.bbclass). # pinned source (forgefirm-image-manifest.bbclass).
SRCREV = "6573abd725e885f0a56463cc83bcb1863525ff77" SRCREV = "66148337ee748ce6f57a398c4ba5e7452319fa07"
PV = "0.1.0" PV = "0.1.0"