A field install failed on "firmware download failed", and worked after a
reboot. The download was one bare curl -fL: no retry, no resume, no bound
on a stalled transfer, and nothing on the machine recorded what had gone
wrong.
The download. download_fw makes up to five tries, 5, 15, 30 and 60 seconds
apart. Each try resumes the partial file (curl -C -) and is bounded: 20 s
to connect, and a transfer below 1 KB/s for 30 s ends the try. The file is
written as forgefirm.fw.part and takes its name only when curl finished;
the signature check that follows is what vouches for its content. A full
disk (curl 23) and a release that is not there (HTTP 404) end the tries at
once, because waiting cannot fix them. A partial file the server will not
resume (curl 33 or 36, HTTP 416) starts over. The loop is the installer's
own rather than curl --retry: the factory curl on the bench reference is
7.69.1, whose --retry does not count a resolver failure or a dropped
transfer as retryable, and older factory builds carry older curls. The
owner sees the reason in words with each retry, and the final failure says
that a re-run goes straight to the download, because the archives are kept.
The log. Every run appends to /data/log/forgefirm/install/install.log, in
the log tree's own line format (UTC, program "install"): the installer's
md5 (which revision ran), the factory version and the slots, the owner's
answers, each archive, each download try with curl's exit code, the HTTP
code and the reason, the machine's clock at each try (a wrong clock breaks
TLS), and after a failed try the address, the default route, the resolver
and whether github.com resolves; then the signature and identity checks,
the write, the boot selection, and the reason for any failure through
die(). The log is appended across runs, so the run that failed is still
there after the run that worked. Logging never fails the install.
forgectrl's log export carries the directory (forgectrl 0dae758).
Proven: tests/test_installer.py runs the installer's own functions under
sh against a scripted curl - a clean download, a resolver failure and a
dropped transfer that resume to the full file, the tries running out, 404
and a full disk ending them at once, a stale partial file starting over,
the TLS reason naming the clock, every log line in the tree format, die()
leaving its reason, and an unwritable log not failing the run. The whole
host suite, 409 tests, passes under Linux and the coverage lint is clean.
Bench: the same functions under the factory firmware's own shell (busybox
1.31.1 ash, the factory slot of the bench reference in a chroot) resumed,
retried, ran out of tries and logged exactly as under sh.
Acceptance: logs.tree-tail-export now plants a probe file in the install
directory and requires it back in the export bundle, its line intact and
its MAC and IPv4 address redacted, and requires an install log in the
bundle when the machine has one. The installer itself is not on the image:
the install page fetches it from master, so it is live with this push.
The archive was written straight to its final name and a rerun accepted
any non-empty file as complete, so a run interrupted mid-archive left a
truncated image that the next run kept and then overwrote the slot. The
archive is now written as .part and renamed on success; an existing
archive counts only with its manifest line present and a whole gzip
stream, otherwise it is archived again.
The recipe URLs, the release and install URLs, the vendor check and the CI checkouts name openglow-org, and the grblHAL core fork is openglow-org/grblHAL-core. No catalog change is owed: the recipe edits move the meta-forgefirm content hash, which every test fingerprint folds in through the platform block, so the whole catalog re-runs on its own.
- Controllers stop at K80, before forgectrl at K90: runlevel 0/6 no
longer tears down the cooling engine, fire gates, and broker while a
controller may still be executing a job.
- The grblhal/gfcloud init scripts are real emergency levers: stop
routes through the supervisor (POST /controller/stop - a bare pkill
was safed and respawned seconds later), start resumes supervision,
status exists, and the pkill fallback matches full executable paths
instead of truncated names or bare substrings.
- slotmigrate: the partition grow gets the same 2048-sector tolerance
as the filesystem branch (an exact compare rewrote the MBR at S02 on
every boot on disks where the grow cannot land on the last sector),
verifies it made progress, and the resize2fs retry is bounded at
three attempts with the counter kept on p3 itself.
- Installer: archive product/platform are verified after the signature,
and a validly signed OLDER release now requires an explicit yes
instead of installing as a silent downgrade. All predictable /tmp
paths in the installer and ffboot are mktemp now.
- release.sh rejects multiple positional versions (the last one used to
win silently) and a release without factory-era verification dies
unless explicitly bypassed; mkfw.sh refuses to pack when the public
key for the post-sign self-check is missing.
- forgefirm-logrotate: size-capped rotation (boot + hourly) for the
/data logs - a full /data breaks settings, update staging, and the
controllers own writes.
- Bench build scripts derive every path from their own location or
FF_SRC_TOP/FF_BUILD_TOP and log to mktemp files.
- Move the passwordless-root debug-tweaks image feature out of the
shared kas config into forgefirm-image-dev.bb, so the release
forgefirm-image built from the same config is not passwordless-root.
release.sh gains a gate that reads the built rootfs /etc/shadow and
fails on an empty root password, plus a config-level guard that
debug-tweaks is not present in the resolved kas dump. (B-1)
- The installer copies ffboot out of the signature-verified new rootfs
it already mounts, instead of fetching and executing it from a mutable
GitHub raw ref. (B-2)
- Record audit remediation Phase 2 (GATE B) status in BRINGUP.md,
including the bench pass still required to close the gate.
Before writing the target slot, the installer now shows what it holds
(factory firmware v<ver>, ForgeFIRM, an unrecognized filesystem, or
unknown/unreadable content). Factory images are archived as before;
anything else requires the operator to type ERASE, since it is
overwritten without a backup. The archive manifest now records the
semantic FIRMWARE_VERSION (ver=), which the update manager displays in
the restore list. Bump forgectrl to the matching GUI change.
The embedded pubkey is the production key from the signing ceremony;
release.sh's key-match gate now refuses any other signer. Verified:
production-signed archives pass fwup 1.16 and the factory's 0.14.2;
dev-signed archives are rejected.
The image's fstab keeps the factory slots mounted under /factory, and
busybox mount's auto-type iteration against an already-mounted ext4
device provokes a cosmetic kernel 'Can't open blockdev' for each
foreign-type claim (reproduced and pinned on the bench: ext3-typed
mount of an ext4-held device prints it; ext4-typed does not). Probes
now reuse an existing mountpoint from /proc/mounts and mount fresh
targets with an explicit -t ext4.
Newer factory firmware's generic /etc/fw_env.config points at the
wrong device; its per-device /etc/fw_env_mmcblk2.config is the correct
one for the eMMC environment. The read-back verify caught the failed
write and aborted before the flip, as designed.
dd|gzip runs backgrounded while the installer prints compressed MB
every few seconds (old busybox dd has no status=progress); dd's exit
status is captured through a file so a device read failure is not
masked by gzip succeeding on truncated input.
Newer factory firmware (2024) has no /factory/imgN mounts and a
read-only rootfs, so slot probing and post-write verification mount
under /tmp, with the active slot read from the running root. The
target-slot unmount sweeps /proc/mounts (older firmware DOES mount the
slots). A failed ffboot download keeps an existing /data/ffboot
instead of aborting, so a local-.fw install works fully offline.
install-forgefirm.sh is now single-stage and never repartitions: run
from factory firmware, it archives every factory slot version plus the
recovery boot partitions to /data/forgefirm/archive (manifest with
md5s), verifies the signed forgefirm.fw against the embedded ForgeFIRM
pubkey (raw 32-byte form for the factory's fwup 0.14.2; dev key until
the production key ceremony), applies it to the INACTIVE slot with the
factory's own fwup, post-verifies the written rootfs, installs
/data/ffboot, and flips the saved env with read-back verification. The
booted factory slot stays installed and bootable; /data is untouched
beyond the archive. Fixed release asset name forgefirm.fw (version in
the fwup metadata and release tag).
slotmigrate (new recipe, rcS before mountall) reclaims the legacy
layout on eMMC-slot boots: deletes p4, grows p3 to the end of the
disk (sfdisk + partx BLKPG - works with a sibling partition as root),
then e2fsck+resize2fs. Every step is keyed off the actual disk state,
so interrupted runs resume and factory-layout disks are a no-op; SD
boots never touch the eMMC.
install-forgefirm.sh completed silently broken when the download, flash
write, mount or uEnv rewrite failed - add die() checks around every
critical step (audit N13/M12) and download the release asset under the
exact Scarthgap artifact name (forgefirm-image-glowforge.rootfs.wic.gz)
so uploads need no renaming. ffboot no longer depends on the never-
provisioned /etc/fw_env_mmcblk2.config: it falls back to
/etc/fw_env.config (which both the factory and ForgeFIRM images ship,
pointing at the eMMC env) and checks fw_setenv results (audit N16).
BUILD.md gets the real artifact name and marks the built u-boot
reference-only; INSTALL.md drops the stale script/ogboot names and the
cloud-connect promise, stating the image is bring-up-only (audit N11).