mirror of
https://github.com/openglow-org/forgefirm.git
synced 2026-09-27 16:51:12 -07:00
forgectrl pinned at the bench-proven GPU path; the session's record
The pin moves to forgectrl 6614833: the five hardware corrections from the GPU demosaic's first bench session (surfaceless EGL config, ARGB render targets, GR88 raw import, the ipu_copy stride-fix crop between the GPU and the CODA, the chroma mirror width). CAMPAIGN-LOG carries the dated record of how each was found; BRINGUP item 20 now separates what that session proved (the full GPU -> IPU -> VPU path serving correct frames, H.264 as valid fragmented MP4 on hardware, texture limits, slot fit) from what remains (the 140 ms render, the bottom row, MSE playback, the CPU measurement, frame skip, coexistence, the campaign).
This commit is contained in:
+26
-23
@@ -1446,29 +1446,32 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`.
|
|||||||
are established. A pause on contact, on the factory's shape, would be
|
are established. A pause on contact, on the factory's shape, would be
|
||||||
the first use.
|
the first use.
|
||||||
|
|
||||||
20. **Video pipeline offload — bench validation.** Code-complete and
|
20. **Video pipeline offload — bench validation.** First hardware session
|
||||||
host-proven (CI: `mp4mux_test`; catalog: `camera.h264-stream`), not yet
|
done (2026-08-24, dev image 20260824122014, drill binaries; fixes in
|
||||||
run on hardware: the GC880 GLES2 demosaic (`src/gpu_debayer.c`,
|
forgectrl 6614833). Proven: Mesa etnaviv fits the release slot
|
||||||
dlopen'd Mesa, capture dmabuf in, encoder dmabuf out), the CODA960
|
(~5 MiB margin); surfaceless EGL and dmabuf import both directions;
|
||||||
H.264 stream (`/cam/h264`, fragmented MP4, panel MSE player,
|
`GL_MAX_TEXTURE_SIZE` 8192 (no tiling even at 8 MP); the full path
|
||||||
~1.5 Mbit/s vs MJPEG's ~9), and the CSI hardware frame skip behind
|
GPU render → IPU stride-fix crop (`src/ipu_copy.c`, the render
|
||||||
`FORGECTRL_STREAM_FPS`. Falls back to the NEON path wherever a piece
|
engine's 64-byte rows and the CODA's round_up(width,16) stride never
|
||||||
is missing, so the existing stream is not at risk. To settle on the
|
meet, so the IPU crops between them, 14 ms, no CPU touch) → VPU
|
||||||
bench, in order: (1) the image builds with Mesa etnaviv inside the
|
encode, `convert: "gpu"`, image correct to within 2 counts of the
|
||||||
200 MiB slot cap (distro `PACKAGECONFIG:pn-mesa`, image adds
|
scalar demosaic; `/cam/h264` serving valid fragmented MP4 on
|
||||||
`libegl-mesa libgles2-mesa libgbm mesa-megadriver`); (2) surfaceless
|
hardware (avc1.424020, ~480 kbit/s on a static bed). Remaining, in
|
||||||
EGL comes up and dmabuf import works both directions (the open logs
|
order: (a) **GPU render time, 140 ms/frame** (~6 fps ceiling; the
|
||||||
say exactly which probe refused); (3) `FORGECTRL_GPU_CHECK` reports a
|
core runs its full 528 MHz, suspicion is the pre-HALTI
|
||||||
max delta of a couple of counts against the scalar demosaic;
|
linear-texture sampling path - decompose per-pass, consider a tiled
|
||||||
(4) `GL_MAX_TEXTURE_SIZE` covers 1944 rows (5 MP single-tile; the
|
shadow copy or fewer chroma taps); (b) the **bottom output row**
|
||||||
8 MP tiling path can only be exercised on an HD machine or by
|
differs from the CPU demosaic (1296 samples, max delta 134; suspect
|
||||||
forcing tiles); (5) coda H.264 rate control and picture quality at
|
the IC's last-line behavior); (c) browser MSE playback of the panel's
|
||||||
1296x972p15, and the measured CPU with a panel H.264 viewer vs the
|
H.264 view; (d) measured CPU with an H.264 viewer vs the 41 % MJPEG
|
||||||
41 % MJPEG baseline; (6) the stream-during-jog coexistence drill
|
baseline; (e) CSI hardware frame-skip exercise; (f) the
|
||||||
rerun with the GPU path active. Switches to strip a suspect layer:
|
stream-during-jog coexistence drill with the GPU path active;
|
||||||
`FORGECTRL_NO_GPU`, `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`,
|
(g) `camera.h264-stream` and the full campaign on an image carrying
|
||||||
plus the existing `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` /
|
the fixes. Switches to strip a suspect layer: `FORGECTRL_NO_GPU`,
|
||||||
`FORGECTRL_NO_CACHED_BUFS`.
|
`FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, plus the existing
|
||||||
|
`FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` /
|
||||||
|
`FORGECTRL_NO_CACHED_BUFS`; `FORGECTRL_GPU_CHECK` tightens the
|
||||||
|
stream-stats cadence and logs the render-versus-copy split.
|
||||||
|
|
||||||
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the
|
**Deliberately not gated:** an armed GRBL job after an underrun cuts at the
|
||||||
stale origin unless homing is required (GRBL mode permits unhomed cutting; the
|
stale origin unless homing is required (GRBL mode permits unhomed cutting; the
|
||||||
|
|||||||
@@ -3630,6 +3630,45 @@ on the bed, and the two together by one real print. Still open from that
|
|||||||
plan: the bench actuator for the lid, interlock and button, and the
|
plan: the bench actuator for the lid, interlock and button, and the
|
||||||
finer coverage maps.
|
finer coverage maps.
|
||||||
|
|
||||||
|
## 2026-08-24: the GPU demosaic's first light, and what the probes caught
|
||||||
|
|
||||||
|
Dev image 20260824122014 (the first with Mesa etnaviv; release ext4
|
||||||
|
204,140,544 bytes, ~5 MiB under the slot cap) flashed by the operator;
|
||||||
|
the drill ran forgectrl builds from `/tmp` against the running image,
|
||||||
|
each iteration probed over the stream, `FORGECTRL_GPU_CHECK`, and frame
|
||||||
|
captures diffed against the CPU demosaic on the host.
|
||||||
|
|
||||||
|
Five faults found and fixed in one session, each named by a probe log
|
||||||
|
line or the compare (forgectrl 6614833):
|
||||||
|
|
||||||
|
1. `eglChooseConfig` returned nothing: EGL_SURFACE_TYPE defaults to
|
||||||
|
WINDOW_BIT and the surfaceless platform has no window configs. Ask
|
||||||
|
for surface type 0.
|
||||||
|
2. Every fourth output byte was 255: the render engine writes an
|
||||||
|
XRGB8888 surface's undefined X byte as opaque. Render ARGB8888.
|
||||||
|
3. The raw import failed etnaviv's stride check (width padded to 16
|
||||||
|
texels): 2592 bytes is 1296 GR88 texels exactly, not 656 padded
|
||||||
|
XRGB ones. Import GR88, one texel per Bayer pair.
|
||||||
|
4. The GPU cannot write the CODA's buffers at all: 64-byte render rows
|
||||||
|
versus round_up(width,16) strides never meet at these widths. New
|
||||||
|
`ipu_copy` module: render into the IPU CSC/scaler's wider source
|
||||||
|
(stride align(w,128)) and let the IPU crop into each encoder over
|
||||||
|
dmabuf. 14 ms a copy, no CPU touch.
|
||||||
|
5. The chroma mirror used a quarter-width plane where the plane is
|
||||||
|
half-width: the right half of both chroma planes clamped to
|
||||||
|
column 0. The three-frame diff-by-transform analysis on the host
|
||||||
|
named both this and fault 2.
|
||||||
|
|
||||||
|
End state on the bench: `convert: "gpu"` serving MJPEG, the GPU/CPU
|
||||||
|
compare clean to 2 counts except the bottom row (1296 samples, max
|
||||||
|
delta 134, unexplained); `/cam/h264` delivering valid fragmented MP4
|
||||||
|
from the CODA BIT processor (avc1.424020, ~480 kbit/s on the static
|
||||||
|
bed). Open, measured: the render costs 140 ms a frame against the IPU's
|
||||||
|
14 (GPU at its full 528 MHz - the suspicion is pre-HALTI linear-texture
|
||||||
|
sampling), so the GPU path holds ~6 fps until that is run down. The
|
||||||
|
`getenv` implicit-declaration fix in debayer.c rode along. Bench left
|
||||||
|
clean; stock service restored.
|
||||||
|
|
||||||
## Superseded status notes
|
## Superseded status notes
|
||||||
|
|
||||||
### Shared machine services — remaining polish, as listed 2026-08-13
|
### Shared machine services — remaining polish, as listed 2026-08-13
|
||||||
|
|||||||
@@ -2,5 +2,5 @@
|
|||||||
# only SRCREV and PV here - the image manifest leaves *-pin.inc out of the
|
# only SRCREV and PV here - the image manifest leaves *-pin.inc out of the
|
||||||
# layer content hash because the component entry already identifies the
|
# layer content hash because the component entry already identifies the
|
||||||
# pinned source (forgefirm-image-manifest.bbclass).
|
# pinned source (forgefirm-image-manifest.bbclass).
|
||||||
SRCREV = "6573abd725e885f0a56463cc83bcb1863525ff77"
|
SRCREV = "66148337ee748ce6f57a398c4ba5e7452319fa07"
|
||||||
PV = "0.1.0"
|
PV = "0.1.0"
|
||||||
|
|||||||
Reference in New Issue
Block a user