diff --git a/docs/BRINGUP.md b/docs/BRINGUP.md index 32d6938..f10f929 100644 --- a/docs/BRINGUP.md +++ b/docs/BRINGUP.md @@ -484,6 +484,10 @@ bounce copy; daemon ~41 % CPU with one viewer. Full-res snapshot 2.4 s warm / copy, needed on a kernel without the `allow_cache_hints` patch — detected via the `MMAP_CACHE_HINTS` capability bit), `FORGECTRL_NO_NEON` (scalar convert, bit-identical), `FORGECTRL_NO_VPU` (libjpeg; also the snapshot path). +Newer and **not yet bench-run** (Next work item 20): the GC880 GPU demosaic, +the `/cam/h264` CODA960 H.264 stream, and CSI hardware frame skip, with +`FORGECTRL_NO_GPU` / `FORGECTRL_NO_H264` / `FORGECTRL_NO_HW_SKIP` to strip +them individually; all fall back to the measured paths above. `/cam/status` reports `encoder` and `buffers`. A CSI glitch frame can out-size the coda driver's default JPEG capture buffer, so forgectrl requests 3 B/px and drops error-flagged dequeues as single bad frames. **LightBurn consumes the @@ -1442,6 +1446,30 @@ Open items only. Anything closed is in `CAMPAIGN-LOG.md`. are established. A pause on contact, on the factory's shape, would be the first use. +20. **Video pipeline offload — bench validation.** Code-complete and + host-proven (CI: `mp4mux_test`; catalog: `camera.h264-stream`), not yet + run on hardware: the GC880 GLES2 demosaic (`src/gpu_debayer.c`, + dlopen'd Mesa, capture dmabuf in, encoder dmabuf out), the CODA960 + H.264 stream (`/cam/h264`, fragmented MP4, panel MSE player, + ~1.5 Mbit/s vs MJPEG's ~9), and the CSI hardware frame skip behind + `FORGECTRL_STREAM_FPS`. Falls back to the NEON path wherever a piece + is missing, so the existing stream is not at risk. To settle on the + bench, in order: (1) the image builds with Mesa etnaviv inside the + 200 MiB slot cap (distro `PACKAGECONFIG:pn-mesa`, image adds + `libegl-mesa libgles2-mesa libgbm mesa-megadriver`); (2) surfaceless + EGL comes up and dmabuf import works both directions (the open logs + say exactly which probe refused); (3) `FORGECTRL_GPU_CHECK` reports a + max delta of a couple of counts against the scalar demosaic; + (4) `GL_MAX_TEXTURE_SIZE` covers 1944 rows (5 MP single-tile; the + 8 MP tiling path can only be exercised on an HD machine or by + forcing tiles); (5) coda H.264 rate control and picture quality at + 1296x972p15, and the measured CPU with a panel H.264 viewer vs the + 41 % MJPEG baseline; (6) the stream-during-jog coexistence drill + rerun with the GPU path active. Switches to strip a suspect layer: + `FORGECTRL_NO_GPU`, `FORGECTRL_NO_H264`, `FORGECTRL_NO_HW_SKIP`, + plus the existing `FORGECTRL_NO_VPU` / `FORGECTRL_NO_NEON` / + `FORGECTRL_NO_CACHED_BUFS`. + **Deliberately not gated:** an armed GRBL job after an underrun cuts at the stale origin unless homing is required (GRBL mode permits unhomed cutting; the underrun itself alarms and unlinks the anchor). Not in the acceptance catalog diff --git a/docs/VIDEO.md b/docs/VIDEO.md index 63c0a25..a3049d2 100644 --- a/docs/VIDEO.md +++ b/docs/VIDEO.md @@ -1,10 +1,11 @@ # Video and the cameras The machine has two cameras — one in the lid looking down at the bed, one in the -print head looking at the material under the lens — and ForgeFIRM serves both as -plain **MJPEG over HTTP** from the web control panel. There is no app, no cloud -relay, and no proprietary protocol: a browser, LightBurn, or anything else that -can read an MJPEG stream or fetch a JPEG can use them. +print head looking at the material under the lens — and ForgeFIRM serves both +over plain HTTP from the web control panel: **MJPEG** for anything that can read +a stream of JPEGs, and an **H.264** live stream for clients that decode video +(the panel uses it when the browser can). There is no app, no cloud relay, and +no proprietary protocol. One rule governs all of it: **the cameras only capture with the lid closed** (§2). Everything else here assumes that condition is met. @@ -115,6 +116,7 @@ buttons; **Live** switches the same frame to the running stream. |---|---| | `/cam/stream?cam=lid` | continuous MJPEG (`multipart/x-mixed-replace`) | | `/cam/stream?cam=head` | the same, from the head camera | +| `/cam/h264?cam=lid` | continuous H.264 as fragmented MP4 (see §5.6): the same picture in a fraction of the bytes, for clients that decode video | | `/cam/snapshot?cam=lid` | one full-resolution JPEG | | `/cam/snapshot?cam=lid&res=half` | one half-resolution JPEG (much faster) | | `/cam/status` | JSON: which sensor, which camera, frame rate, frame sizes, whether the lid currently permits capture | @@ -145,7 +147,7 @@ practice: paste the URL into any local client and it works. | Live stream | 1296 × 972 | 1632 × 1224 | | Full snapshot | 2592 × 1944 | 3264 × 2448 | | Half snapshot | 1296 × 972 | 1632 × 1224 | -| Format | JPEG, quality 75 by default | same | +| Stream formats | MJPEG (quality 75 by default) and H.264 (~1.5 Mbit/s by default) | same | | Frame rate | **15 fps** sustained | not yet measured (§10) | Measured on a 5 MP machine: 15.0 fps with a viewer attached, which is the rate @@ -161,6 +163,11 @@ sensor pixels becomes exactly one output pixel, which is why the stream is precisely half the capture in each axis and why it is cheap enough to run continuously. +The 41 % figure is the NEON demosaic feeding the hardware JPEG encoder. When +the GPU demosaic and the H.264 stream carry the load instead (§5.6), the +stream's CPU cost drops to bookkeeping; those two paths are newer than the +figure above and their own numbers will be measured on the bench the same way. + --- ## 5. What the sensor can do versus what ForgeFIRM sends @@ -177,7 +184,7 @@ reason. | Bit depth | 10 bits per pixel | 8 bits | JPEG is 8-bit, and 8-bit is what makes §5.2 fit | | Exposure / color | auto exposure and auto white balance | fixed values | a bed image has to look the same frame to frame; §5.4 | | Lens | — | no correction applied | correction belongs in the client; §5.5 | -| Encoding | — | MJPEG only, nothing recorded | §5.6 | +| Encoding | — | MJPEG and H.264, nothing recorded | §5.6 | | Mirroring | a mirror register | mirrored in software instead | the register breaks capture on this board; §5.7 | ### 5.1 Resolution and frame rate on a 5 MP machine @@ -273,20 +280,43 @@ correction on the host, which is more accurate than a fixed correction baked into the firmware and costs the machine's CPU nothing. Run that calibration before trusting the camera overlay for placement. -### 5.6 MJPEG only — and nothing is recorded +### 5.6 Two streams, one picture: MJPEG and H.264. Nothing is recorded. -The stream is a sequence of complete JPEG frames, not H.264 or any other -inter-frame codec, and **the machine never writes video to disk**. +The same live picture is served two ways, and **the machine never writes video +to disk**. -MJPEG is the right trade here: every frame stands alone, so a viewer can join -or leave at any moment and a dropped frame costs nothing; browsers and sender -software consume it with no plugin; and stream frames are encoded by the -board's **hardware JPEG encoder**, which is what makes 15 fps affordable while -the machine is also running a job. (Stills are encoded in software instead, -which is most of why a full-resolution one takes a couple of seconds.) An -inter-frame codec would need buffering and a container, would break the "any -client, any time" property, and would buy bandwidth savings that a LAN does -not need. +**MJPEG** (`/cam/stream`) is the universal one: every frame is a complete +JPEG, so a viewer can join or leave at any moment, a dropped frame costs +nothing, and browsers, LightBurn and mjpg-streamer clients consume it with no +plugin. It stays, unchanged, and it is what anything that cannot decode video +should use. + +**H.264** (`/cam/h264`) exists because bytes on this machine are not free. +MJPEG re-sends the whole scene fifteen times a second, roughly 9 Mbit/s, and +the WiFi transmit path runs on the machine's single CPU core, where measured +cost is about 7 % of the core per MB/s sent. A bed camera's scene barely +changes between frames, which is exactly what an inter-frame codec exploits: +the H.264 stream carries the same picture in roughly 1.5 Mbit/s and gives most +of that CPU back. It arrives as fragmented MP4, the form a browser's Media +Source Extensions accept, with the codec named in an `X-H264-Codec` response +header; the panel's **Live** button uses it automatically where the browser +can and falls back to MJPEG where it cannot. Latency is a beat behind MJPEG +(under a second), which is why LightBurn keeps consuming the MJPEG stream. + +Both encoders are hardware: JPEG frames come from the CODA960's JPEG unit and +H.264 from its BIT processor, two independent engines, so serving both at once +does not double any cost that matters. The demosaic that feeds them runs as +fragment shaders on the SoC's GC880 GPU when the image ships the GL stack +(reported as `"convert": "gpu"` in `/cam/status`), reading the sensor frame +and writing the encoder's buffer directly, so a stream frame never crosses the +CPU at all; without the GPU it falls back to the NEON demosaic. (Stills are +still demosaiced and encoded on the CPU, which is most of why a +full-resolution one takes a couple of seconds.) + +One more consumer of nothing: with a frame-rate cap set (`FORGECTRL_STREAM_FPS` +of 1 or more), the cap is programmed into the CSI receiver's frame-skip +hardware, and skipped frames are dropped before they are ever written to +memory. `/cam/status` reports `"hw_fps_skip": true` when that is in effect. If you want a recording, record the stream on the computer watching it. The machine stores its firmware, settings and logs on a small internal flash device diff --git a/forgetest/forgetest/suite/camera.py b/forgetest/forgetest/suite/camera.py index d31962a..a2a7551 100644 --- a/forgetest/forgetest/suite/camera.py +++ b/forgetest/forgetest/suite/camera.py @@ -5,6 +5,8 @@ from ..baseline import LID_LAMP_ATTR _CAM_COVERS = [("forgectrl", "src/cam.*"), ("forgectrl", "src/camhealth.*"), ("forgectrl", "src/debayer.*"), ("forgectrl", "src/vpu_jpeg.*"), + ("forgectrl", "src/vpu_h264.*"), ("forgectrl", "src/mp4mux.*"), + ("forgectrl", "src/gpu_debayer.*"), ("forgectrl", "src/main.c"), ("python3-gfhardware", "gfhardware/src/**"), ("python3-gfhardware", "gfhardware/cam*")] @@ -158,6 +160,68 @@ def sensor_profile(ctx): ctx.check(got == (w, h), "%s snapshot is %s, not %dx%d", res, got, w, h) +@test("camera.h264-stream", title="H.264 live stream", subsystem="camera", + kind="auto", est_min=2, + covers=_CAM_COVERS, requires=["forgectrl.panel-serves"], + steps=["Setup: lid closed."], + description="/cam/h264 answers a fragmented MP4: a codec header naming the SPS profile, an " + "init segment (ftyp+moov) followed by media fragments (moof+mdat) within a few " + "seconds, and /cam/status reporting the H.264 encoder active with this client " + "counted. Also records which demosaic path (GPU or CPU) served the stream. On a " + "machine whose image lacks Mesa or whose encoder refused, the endpoint must " + "answer 503 rather than hang - that is a pass for the endpoint but is recorded " + "in the evidence.") +def h264_stream(ctx): + import urllib.error + import urllib.request + + fc = ctx.forgectrl + ev = ctx.evidence + + req = urllib.request.Request(fc.base + "/cam/h264?cam=lid", + headers={"Host": fc.host_header()}) + try: + r = urllib.request.urlopen(req, timeout=20) + except urllib.error.HTTPError as e: + ev["status"] = e.code + ev["body"] = e.read()[:200].decode("utf-8", "replace") + ctx.check(e.code == 503, "H.264 refused with %s, not 503: %s", e.code, ev["body"]) + ctx.log("H.264 unavailable on this machine (503): %s", ev["body"]) + return + with r: + codec = r.headers.get("X-H264-Codec", "") + ctype = r.headers.get("Content-Type", "") + ev["codec"] = codec + ev["content_type"] = ctype + ctx.log("codec %s, content-type %s", codec, ctype) + ctx.check(ctype == "video/mp4", "content type is %r", ctype) + ctx.check(codec.startswith("avc1."), "codec header is %r", codec) + + # Init segment, then at least one media fragment. + data = b"" + while len(data) < 512 * 1024: + chunk = r.read(65536) + if not chunk: + break + data += chunk + if b"moof" in data and b"mdat" in data: + break + ev["bytes"] = len(data) + ctx.check(data[4:8] == b"ftyp", "stream does not start with ftyp") + ctx.check(b"moov" in data, "no moov (init segment)") + ctx.check(b"avcC" in data, "no avcC in the init segment") + ctx.check(b"moof" in data and b"mdat" in data, + "no media fragment arrived in %d bytes", len(data)) + + st, body = fc.get("/cam/status") + ctx.check(st == 200 and isinstance(body, dict), "GET /cam/status -> %s", st) + ev["cam_status"] = body + ctx.log("convert=%s h264=%s", body.get("convert"), body.get("h264")) + h = body.get("h264") or {} + ctx.check(h.get("active") is True, "/cam/status does not report the H.264 encoder active: %s", + body) + + @test("camera.frame-health", title="Capture delivers whole frames", subsystem="camera", kind="auto", est_min=1, covers=_CAM_COVERS, requires=["forgectrl.panel-serves"], diff --git a/meta-forgefirm/conf/distro/forgefirm.conf b/meta-forgefirm/conf/distro/forgefirm.conf index b5e7b92..32e64a2 100644 --- a/meta-forgefirm/conf/distro/forgefirm.conf +++ b/meta-forgefirm/conf/distro/forgefirm.conf @@ -5,9 +5,17 @@ DISTRO_NAME = "OpenGlow/ForgeFIRM" DISTRO_VERSION = "0.0.0" DISTRO_FEATURES:remove = " \ - 3g alsa avahi bluetooth bluez5 ext2 ipv6 irda nfc nfs opengl pci pcmcia \ + 3g alsa avahi bluetooth bluez5 ext2 ipv6 irda nfc nfs pci pcmcia \ pulseaudio vulkan wayland x11 zeroconf " +# opengl stays: forgectrl's camera demosaic runs as GLES2 fragment +# shaders on the GC880 (etnaviv), reached through surfaceless EGL with +# no display stack. Mesa is trimmed to exactly that: the etnaviv +# gallium driver, GLES/EGL/GBM, no GLX, no X11/Wayland platforms +# (both remain removed above), so the cost is the libraries and the one +# driver, not a graphics stack. +PACKAGECONFIG:pn-mesa = "opengl gles egl gbm gallium etnaviv" + # System logger: rsyslog replaces busybox syslogd/klogd. Every ForgeFIRM # process logs through it and it is the only log writer (one directory # per logger under /data/log/forgefirm; per-logger levels and the remote diff --git a/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc b/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc index bd55c96..1bb6361 100644 --- a/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc +++ b/meta-forgefirm/recipes-forgefirm/forgectrl/forgectrl-pin.inc @@ -2,5 +2,5 @@ # only SRCREV and PV here - the image manifest leaves *-pin.inc out of the # layer content hash because the component entry already identifies the # pinned source (forgefirm-image-manifest.bbclass). -SRCREV = "3acd66425dd53b3a8b16ae175ec442a4d057b9ac" +SRCREV = "6573abd725e885f0a56463cc83bcb1863525ff77" PV = "0.1.0" diff --git a/meta-forgefirm/recipes-forgefirm/images/forgefirm-image.bb b/meta-forgefirm/recipes-forgefirm/images/forgefirm-image.bb index 2154c5e..94d59d3 100644 --- a/meta-forgefirm/recipes-forgefirm/images/forgefirm-image.bb +++ b/meta-forgefirm/recipes-forgefirm/images/forgefirm-image.bb @@ -35,6 +35,11 @@ IMAGE_INSTALL:remove = "gfui-client" # VIRTUAL-RUNTIME_base-utils-syslog (conf/distro/forgefirm.conf). IMAGE_INSTALL:append = " grblhal-glowforge forgectrl gfhome gfcloud v4l-utils fwup ffboot slotmigrate forgefirm-logging" +# Mesa GLES2/EGL on etnaviv for forgectrl's GPU demosaic (loaded with +# dlopen at runtime; forgectrl itself has no build-time GL dependency, +# and without these packages it falls back to the NEON path). +IMAGE_INSTALL:append = " libegl-mesa libgles2-mesa libgbm mesa-megadriver" + # NXP's firmware EULA covers the i.MX VPU/EPDC blobs the BSP installs, so the # image ships the license text with them (/usr/share/licenses/firmware-imx). # The SDMA firmware brings its own -license package through linux-firmware.