mirror of
https://github.com/openglow-org/forgefirm.git
synced 2026-09-28 01:01:12 -07:00
forgectrl pinned at the bench-proven GPU path; the session's record
The pin moves to forgectrl 6614833: the five hardware corrections from the GPU demosaic's first bench session (surfaceless EGL config, ARGB render targets, GR88 raw import, the ipu_copy stride-fix crop between the GPU and the CODA, the chroma mirror width). CAMPAIGN-LOG carries the dated record of how each was found; BRINGUP item 20 now separates what that session proved (the full GPU -> IPU -> VPU path serving correct frames, H.264 as valid fragmented MP4 on hardware, texture limits, slot fit) from what remains (the 140 ms render, the bottom row, MSE playback, the CPU measurement, frame skip, coexistence, the campaign).
This commit is contained in:
+26
-23
@@ -1446,29 +1446,32 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
|
||||
are established. A pause on contact, on the factory's shape, would be
|
||||
the first use.
|
||||
|
||||
20. **Video pipeline offload — bench validation.** Code-complete and
|
||||
host-proven (CI: `mp4mux_test`; catalog: `camera.h264-stream`), not yet
|
||||
run on hardware: the GC880 GLES2 demosaic (`src/gpu_debayer.c`,
|
||||
dlopen'd Mesa, capture dmabuf in, encoder dmabuf out), the CODA960
|
||||
H.264 stream (`/cam/h264`, fragmented MP4, panel MSE player,
|
||||
~1.5 Mbit/s vs MJPEG's ~9), and the CSI hardware frame skip behind
|
||||
`FORGECTRL_STREAM_FPS`. Falls back to the NEON path wherever a piece
|
||||
is missing, so the existing stream is not at risk. To settle on the
|
||||
bench, in order: (1) the image builds with Mesa etnaviv inside the
|
||||
200 MiB slot cap (distro `PACKAGECONFIG:pn-mesa`, image adds
|
||||
`libegl-mesa libgles2-mesa libgbm mesa-megadriver`); (2) surfaceless
|
||||
EGL comes up and dmabuf import works both directions (the open logs
|
||||
say exactly which probe refused); (3) `FORGECTRL_GPU_CHECK` reports a
|
||||
max delta of a couple of counts against the scalar demosaic;
|
||||
(4) `GL_MAX_TEXTURE_SIZE` covers 1944 rows (5 MP single-tile; the
|
||||
8 MP tiling path can only be exercised on an HD machine or by
|
||||
forcing tiles); (5) coda H.264 rate control and picture quality at
|
||||
1296x972p15, and the measured CPU with a panel H.264 viewer vs the
|
||||
41 % MJPEG baseline; (6) the stream-during-jog coexistence drill
|
||||
rerun with the GPU path active. Switches to strip a suspect layer:
|
||||
`FORGECTRL_NO_GPU`, `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`,
|
||||
plus the existing `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` /
|
||||
`FORGECTRL_NO_CACHED_BUFS`.
|
||||
20. **Video pipeline offload — bench validation.** First hardware session
|
||||
done (2026-08-24, dev image 20260824122014, drill binaries; fixes in
|
||||
forgectrl 6614833). Proven: Mesa etnaviv fits the release slot
|
||||
(~5 MiB margin); surfaceless EGL and dmabuf import both directions;
|
||||
`GL_MAX_TEXTURE_SIZE` 8192 (no tiling even at 8 MP); the full path
|
||||
GPU render → IPU stride-fix crop (`src/ipu_copy.c`, the render
|
||||
engine's 64-byte rows and the CODA's round_up(width,16) stride never
|
||||
meet, so the IPU crops between them, 14 ms, no CPU touch) → VPU
|
||||
encode, `convert: "gpu"`, image correct to within 2 counts of the
|
||||
scalar demosaic; `/cam/h264` serving valid fragmented MP4 on
|
||||
hardware (avc1.424020, ~480 kbit/s on a static bed). Remaining, in
|
||||
order: (a) **GPU render time, 140 ms/frame** (~6 fps ceiling; the
|
||||
core runs its full 528 MHz, suspicion is the pre-HALTI
|
||||
linear-texture sampling path - decompose per-pass, consider a tiled
|
||||
shadow copy or fewer chroma taps); (b) the **bottom output row**
|
||||
differs from the CPU demosaic (1296 samples, max delta 134; suspect
|
||||
the IC's last-line behavior); (c) browser MSE playback of the panel's
|
||||
H.264 view; (d) measured CPU with an H.264 viewer vs the 41 % MJPEG
|
||||
baseline; (e) CSI hardware frame-skip exercise; (f) the
|
||||
stream-during-jog coexistence drill with the GPU path active;
|
||||
(g) `camera.h264-stream` and the full campaign on an image carrying
|
||||
the fixes. Switches to strip a suspect layer: `FORGECTRL_NO_GPU`,
|
||||
`FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, plus the existing
|
||||
`FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` /
|
||||
`FORGECTRL_NO_CACHED_BUFS`; `FORGECTRL_GPU_CHECK` tightens the
|
||||
stream-stats cadence and logs the render-versus-copy split.
|
||||
|
||||
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the
|
||||
stale origin unless homing is required (GRBL mode permits unhomed cutting; the
|
||||
|
||||
@@ -3630,6 +3630,45 @@ on the bed, and the two together by one real print. Still open from that
|
||||
plan: the bench actuator for the lid, interlock and button, and the
|
||||
finer coverage maps.
|
||||
|
||||
## 2026-08-24: the GPU demosaic's first light, and what the probes caught
|
||||
|
||||
Dev image 20260824122014 (the first with Mesa etnaviv; release ext4
|
||||
204,140,544 bytes, ~5 MiB under the slot cap) flashed by the operator;
|
||||
the drill ran forgectrl builds from `/tmp` against the running image,
|
||||
each iteration probed over the stream, `FORGECTRL_GPU_CHECK`, and frame
|
||||
captures diffed against the CPU demosaic on the host.
|
||||
|
||||
Five faults found and fixed in one session, each named by a probe log
|
||||
line or the compare (forgectrl 6614833):
|
||||
|
||||
1. `eglChooseConfig` returned nothing: EGL_SURFACE_TYPE defaults to
|
||||
WINDOW_BIT and the surfaceless platform has no window configs. Ask
|
||||
for surface type 0.
|
||||
2. Every fourth output byte was 255: the render engine writes an
|
||||
XRGB8888 surface's undefined X byte as opaque. Render ARGB8888.
|
||||
3. The raw import failed etnaviv's stride check (width padded to 16
|
||||
texels): 2592 bytes is 1296 GR88 texels exactly, not 656 padded
|
||||
XRGB ones. Import GR88, one texel per Bayer pair.
|
||||
4. The GPU cannot write the CODA's buffers at all: 64-byte render rows
|
||||
versus round_up(width,16) strides never meet at these widths. New
|
||||
`ipu_copy` module: render into the IPU CSC/scaler's wider source
|
||||
(stride align(w,128)) and let the IPU crop into each encoder over
|
||||
dmabuf. 14 ms a copy, no CPU touch.
|
||||
5. The chroma mirror used a quarter-width plane where the plane is
|
||||
half-width: the right half of both chroma planes clamped to
|
||||
column 0. The three-frame diff-by-transform analysis on the host
|
||||
named both this and fault 2.
|
||||
|
||||
End state on the bench: `convert: "gpu"` serving MJPEG, the GPU/CPU
|
||||
compare clean to 2 counts except the bottom row (1296 samples, max
|
||||
delta 134, unexplained); `/cam/h264` delivering valid fragmented MP4
|
||||
from the CODA BIT processor (avc1.424020, ~480 kbit/s on the static
|
||||
bed). Open, measured: the render costs 140 ms a frame against the IPU's
|
||||
14 (GPU at its full 528 MHz - the suspicion is pre-HALTI linear-texture
|
||||
sampling), so the GPU path holds ~6 fps until that is run down. The
|
||||
`getenv` implicit-declaration fix in debayer.c rode along. Bench left
|
||||
clean; stock service restored.
|
||||
|
||||
## Superseded status notes
|
||||
|
||||
### Shared machine services — remaining polish, as listed 2026-08-13
|
||||
|
||||
Reference in New Issue
Block a user