diff --git a/docs/BRINGUP.md b/docs/BRINGUP.md index f10f929..632d5ac 100644 --- a/docs/BRINGUP.md +++ b/docs/BRINGUP.md @@ -1446,29 +1446,32 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`. are established. A pause on contact, on the factory's shape, would be the first use. -20. **Video pipeline offload — bench validation.** Code-complete and - host-proven (CI: `mp4mux_test`; catalog: `camera.h264-stream`), not yet - run on hardware: the GC880 GLES2 demosaic (`src/gpu_debayer.c`, - dlopen'd Mesa, capture dmabuf in, encoder dmabuf out), the CODA960 - H.264 stream (`/cam/h264`, fragmented MP4, panel MSE player, - ~1.5 Mbit/s vs MJPEG's ~9), and the CSI hardware frame skip behind - `FORGECTRL_STREAM_FPS`. Falls back to the NEON path wherever a piece - is missing, so the existing stream is not at risk. To settle on the - bench, in order: (1) the image builds with Mesa etnaviv inside the - 200 MiB slot cap (distro `PACKAGECONFIG:pn-mesa`, image adds - `libegl-mesa libgles2-mesa libgbm mesa-megadriver`); (2) surfaceless - EGL comes up and dmabuf import works both directions (the open logs - say exactly which probe refused); (3) `FORGECTRL_GPU_CHECK` reports a - max delta of a couple of counts against the scalar demosaic; - (4) `GL_MAX_TEXTURE_SIZE` covers 1944 rows (5 MP single-tile; the - 8 MP tiling path can only be exercised on an HD machine or by - forcing tiles); (5) coda H.264 rate control and picture quality at - 1296x972p15, and the measured CPU with a panel H.264 viewer vs the - 41 % MJPEG baseline; (6) the stream-during-jog coexistence drill - rerun with the GPU path active. Switches to strip a suspect layer: - `FORGECTRL_NO_GPU`, `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, - plus the existing `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` / - `FORGECTRL_NO_CACHED_BUFS`. +20. **Video pipeline offload — bench validation.** First hardware session + done (2026-08-24, dev image 20260824122014, drill binaries; fixes in + forgectrl 6614833). Proven: Mesa etnaviv fits the release slot + (~5 MiB margin); surfaceless EGL and dmabuf import both directions; + `GL_MAX_TEXTURE_SIZE` 8192 (no tiling even at 8 MP); the full path + GPU render → IPU stride-fix crop (`src/ipu_copy.c`, the render + engine's 64-byte rows and the CODA's round_up(width,16) stride never + meet, so the IPU crops between them, 14 ms, no CPU touch) → VPU + encode, `convert: "gpu"`, image correct to within 2 counts of the + scalar demosaic; `/cam/h264` serving valid fragmented MP4 on + hardware (avc1.424020, ~480 kbit/s on a static bed). Remaining, in + order: (a) **GPU render time, 140 ms/frame** (~6 fps ceiling; the + core runs its full 528 MHz, suspicion is the pre-HALTI + linear-texture sampling path - decompose per-pass, consider a tiled + shadow copy or fewer chroma taps); (b) the **bottom output row** + differs from the CPU demosaic (1296 samples, max delta 134; suspect + the IC's last-line behavior); (c) browser MSE playback of the panel's + H.264 view; (d) measured CPU with an H.264 viewer vs the 41 % MJPEG + baseline; (e) CSI hardware frame-skip exercise; (f) the + stream-during-jog coexistence drill with the GPU path active; + (g) `camera.h264-stream` and the full campaign on an image carrying + the fixes. Switches to strip a suspect layer: `FORGECTRL_NO_GPU`, + `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, plus the existing + `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` / + `FORGECTRL_NO_CACHED_BUFS`; `FORGECTRL_GPU_CHECK` tightens the + stream-stats cadence and logs the render-versus-copy split. **Deliberately not gated:** an armed GRBL job after an underrun cuts at the stale origin unless homing is required (GRBL mode permits unhomed cutting; the diff --git a/docs/CAMPAIGN-LOG.md b/docs/CAMPAIGN-LOG.md index fd60ee4..13a6177 100644 --- a/docs/CAMPAIGN-LOG.md +++ b/docs/CAMPAIGN-LOG.md @@ -3630,6 +3630,45 @@ on the bed, and the two together by one real print. Still open from that plan: the bench actuator for the lid, interlock and button, and the finer coverage maps. +## 2026-08-24: the GPU demosaic's first light, and what the probes caught + +Dev image 20260824122014 (the first with Mesa etnaviv; release ext4 +204,140,544 bytes, ~5 MiB under the slot cap) flashed by the operator; +the drill ran forgectrl builds from `/tmp` against the running image, +each iteration probed over the stream, `FORGECTRL_GPU_CHECK`, and frame +captures diffed against the CPU demosaic on the host. + +Five faults found and fixed in one session, each named by a probe log +line or the compare (forgectrl 6614833): + +1. `eglChooseConfig` returned nothing: EGL_SURFACE_TYPE defaults to + WINDOW_BIT and the surfaceless platform has no window configs. Ask + for surface type 0. +2. Every fourth output byte was 255: the render engine writes an + XRGB8888 surface's undefined X byte as opaque. Render ARGB8888. +3. The raw import failed etnaviv's stride check (width padded to 16 + texels): 2592 bytes is 1296 GR88 texels exactly, not 656 padded + XRGB ones. Import GR88, one texel per Bayer pair. +4. The GPU cannot write the CODA's buffers at all: 64-byte render rows + versus round_up(width,16) strides never meet at these widths. New + `ipu_copy` module: render into the IPU CSC/scaler's wider source + (stride align(w,128)) and let the IPU crop into each encoder over + dmabuf. 14 ms a copy, no CPU touch. +5. The chroma mirror used a quarter-width plane where the plane is + half-width: the right half of both chroma planes clamped to + column 0. The three-frame diff-by-transform analysis on the host + named both this and fault 2. + +End state on the bench: `convert: "gpu"` serving MJPEG, the GPU/CPU +compare clean to 2 counts except the bottom row (1296 samples, max +delta 134, unexplained); `/cam/h264` delivering valid fragmented MP4 +from the CODA BIT processor (avc1.424020, ~480 kbit/s on the static +bed). Open, measured: the render costs 140 ms a frame against the IPU's +14 (GPU at its full 528 MHz - the suspicion is pre-HALTI linear-texture +sampling), so the GPU path holds ~6 fps until that is run down. The +`getenv` implicit-declaration fix in debayer.c rode along. Bench left +clean; stock service restored. + ## Superseded status notes ### Shared machine services — remaining polish, as listed 2026-08-13 diff --git a/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc b/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc index 1bb6361..28049ab 100644 --- a/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc +++ b/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc @@ -2,5 +2,5 @@ # only SRCREV and PV here - the image manifest leaves *-pin.inc out of the # layer content hash because the component entry already identifies the # pinned source (forgefirm-image-manifest.bbclass). -SRCREV = "6573abd725e885f0a56463cc83bcb1863525ff77" +SRCREV = "66148337ee748ce6f57a398c4ba5e7452319fa07" PV = "0.1.0"