The extension sandbox platform: accounts, cgroups, and the deny rules

What the image holds ready before any extension package exists, so that
the first one starts inside it.

forgefirm-sandbox (new recipe, on both images):
  - the account pool: ffx0 to ffx31, uid and gid 800 to 831, one group
    each, /nonexistent, /bin/false, locked. Below 1000 on purpose: the
    forgefirm-users render replaces only the accounts from 1000 up, so an
    account reset leaves the pool alone and the read-only rootfs never
    needs an account made at run time. The image's dynamic system ids
    count down from 999 and stop at 997.
  - an rcS script at S30: cgroup v2 mounted at /sys/fs/cgroup, the cpu,
    memory, and pids controllers handed down to /sys/fs/cgroup/ffx, and
    ffx marked idle-class (cpu.idle; the kernel refuses a cpu.weight on
    top of it, so none is written). The firmware's processes stay in the
    root group. `status` reports both halves and exits nonzero when
    either is missing.
  - /etc/forgefirm/ffx.nft, loaded by the same script, before the network
    starts in rc5: table inet ffx, an output-hook filter with policy
    accept that sends uid 800-831 to chain pool; pool looks the uid up in
    the verdict map `allow`, then answers TCP with a reset (a drop would
    leave a connect to time out) and drops the rest (the sender sees
    EPERM), both counted. The map is the one way through: a uid mapped to
    a chain of that package's destinations. Loading the file again
    replaces the table, allowlists included: it fails closed.

nftables comes in as its runtime dependency, trimmed in the distro config
to the binary and its library with JSON output: no interactive shell, no
Python binding. gmp and jansson were on the image; libmnl and libnftnl are
new. The release rootfs goes from 34.5 to 34.1 MiB free.

scripts/sandbox-rules-test.py, and the workflow sandbox-ci that runs it:
the rule file loaded into a network namespace of its own and sent at from
real uids. Root, 799, and 832 are not touched; 800, 815, and 831 are
refused on 127.0.0.1 and ::1 at once, UDP with EPERM, and a receiver hears
nobody from the pool; an allow chain opens one port on one address to one
uid and nothing else; a reload closes it.

exthost.platform (new suite module exthost.py): the platform proven on a
probe process, not read off a config. In a probe group under ffx, as the
last pool uid: held to cpu.max, stopped by cgroup.freeze and running again
after, stopped at pids.max, killed by the group's own OOM at memory.max
while forgectrl keeps its pid. The 32 accounts as the boot's render left
them. Pool uids 800 and 831 refused TCP to forgectrl on loopback (both
ports, IPv4 and IPv6), to the LAN address, and to the Grbl port, at once,
UDP EPERM, with the rules' counters moving by at least the attempts, while
root reaches the same listeners. A root probe under landlock loses /etc
and TCP connects and keeps /usr; a seccomp filter returns EPERM for the
filtered call. The probe group is removed whatever happens.

Proven. The rules test passes with nft 1.0.9, the image's version, and
three controls each fail it: the range one uid short, the TCP reject
turned to accept, the delete-table line removed. The unit suite passes
(422) with no undefined name. Image 20260920211625 carries all of it (read
back from both rootfs images: 32 accounts in passwd, group, and shadow,
S30forgefirm-sandbox, the rule file and the script byte-identical, nft
with its libraries and no Python binding). On the bench reference, that
image: exthost.platform PASS (5.0 percent of the core under a 5 percent
cpu.max, 0 us frozen and 87358 us thawed over 1.5 s each, 5 of 12 forks
then EAGAIN, rc -9 with oom_kill 1 at a 24 MiB memory.max, the counters
[0, 0] to [12, 4], landlock ABI 6), and forgefirm-sandbox status reports
both halves in place.

Acceptance. exthost.platform gates the platform; sandbox-ci gates the rule
file. The recipe, the rules, and the distro option are layer content, in
the platform identity of every fingerprint.
This commit is contained in:
ScottW514
2026-09-20 18:02:31 -04:00
parent 965f7c3ca9
commit d627ee32f1
9 changed files with 880 additions and 0 deletions
@@ -0,0 +1,44 @@
#!/usr/sbin/nft -f
# Copyright 2026 514 LLC d/b/a OpenGlow
# Written by Scott Wiederhold
# https://community.openglow.org
# SPDX-License-Identifier: MIT
#
# The extension sandbox's network rules (forgefirm-sandbox loads them from
# rcS, before the network starts). Every packet sent from a socket that an
# extension account owns (ffx0 to ffx31, uid 800 to 831) is refused: on
# loopback, on the machine's own LAN address, IPv4 and IPv6 alike. That is
# what keeps a package off the Grbl port, off forgectrl's listeners, and off
# the controller's report route, whatever else fails.
#
# The rules sit on the output hook and match the sending socket's uid, so
# they need no connection tracking, and no packet of the firmware's own pays
# for more than one comparison. A refused TCP connect gets a reset, so it
# fails at once instead of timing out; anything else is dropped, which the
# sender sees as EPERM.
#
# The one way through is the allow map: uid -> a chain holding that package's
# declared destinations. Whatever starts a package adds the element and the
# chain, and removes both when the package stops; a chain that accepts
# nothing returns here and the packet is refused. Loading this file again
# replaces the whole table, allowlists included: it fails closed.
table inet ffx
delete table inet ffx
table inet ffx {
map allow {
typeof meta skuid : verdict
}
chain output {
type filter hook output priority filter; policy accept;
meta skuid 800-831 jump pool
}
chain pool {
meta skuid vmap @allow
meta l4proto tcp counter reject with tcp reset
counter drop
}
}
@@ -0,0 +1,94 @@
#!/bin/sh
# Copyright 2026 514 LLC d/b/a OpenGlow
# Written by Scott Wiederhold
# https://community.openglow.org
# SPDX-License-Identifier: MIT
### BEGIN INIT INFO
# Provides: forgefirm-sandbox
# Required-Start: mountkernfs
# Required-Stop:
# Default-Start: S
# Default-Stop:
# Short-Description: ForgeFIRM extension sandbox: the cgroup tree and the deny rules
### END INIT INFO
# Runs at S30 in rcS, before the network starts (rc5). Two jobs, and each
# is told apart in what it prints and in its exit status, because whatever
# starts extension services checks both in the kernel and starts nothing
# when either is missing:
#
# the cgroup v2 tree mounted at /sys/fs/cgroup, the cpu, memory, and pids
# controllers handed down to /sys/fs/cgroup/ffx and
# from there to the group each package gets. The ffx
# group is idle-class against everything outside it:
# a runnable task in the root group always goes first.
# the deny rules table inet ffx from /etc/forgefirm/ffx.nft: a pool
# uid sends nothing, loopback included.
#
# The firmware's own processes stay in the root group, which has no limit
# and no controller file that could give it one.
CG=/sys/fs/cgroup
RULES=/etc/forgefirm/ffx.nft
CONTROLLERS="cpu memory pids"
cgroup_mounted () {
awk -v m="$CG" '$2 == m && $3 == "cgroup2" { f = 1 } END { exit !f }' /proc/mounts
}
# hand_down <group dir>: every controller into the group's subtree_control.
hand_down () {
for c in $CONTROLLERS; do
echo "+$c" > "$1/cgroup.subtree_control" 2>/dev/null || return 1
done
}
start_cgroups () {
cgroup_mounted || mount -t cgroup2 -o nosuid,nodev,noexec cgroup2 "$CG" || return 1
hand_down "$CG" || return 1
mkdir -p "$CG/ffx" || return 1
hand_down "$CG/ffx" || return 1
# An idle group's weight is the scheduler's own lowest; the kernel refuses
# a cpu.weight written on top of it.
echo 1 > "$CG/ffx/cpu.idle" || return 1
}
start_rules () {
/usr/sbin/nft -f "$RULES"
}
status () {
rc=0
if cgroup_mounted && [ -d "$CG/ffx" ]; then
echo "cgroups: $CG/ffx controls: $(cat "$CG/ffx/cgroup.subtree_control")"
else
echo "cgroups: not set up"; rc=1
fi
if /usr/sbin/nft list table inet ffx >/dev/null 2>&1; then
echo "deny rules: loaded"
else
echo "deny rules: NOT loaded"; rc=1
fi
return $rc
}
case "$1" in
start|restart|reload|force-reload)
rc=0
start_cgroups || { echo "forgefirm-sandbox: the cgroup tree could not be set up"; rc=1; }
start_rules || { echo "forgefirm-sandbox: the deny rules did not load"; rc=1; }
exit $rc
;;
stop)
;;
status)
status
exit $?
;;
*)
echo "Usage: $0 {start|stop|restart|status}"
exit 1
;;
esac
exit 0
@@ -0,0 +1,46 @@
# SPDX-License-Identifier: MIT
SUMMARY = "ForgeFIRM extension sandbox: the account pool, the cgroup tree, and the deny rules"
DESCRIPTION = "What the image holds ready before any extension package runs: \
a static pool of system accounts (ffx0 to ffx31, uid and gid 800 to 831), \
the cgroup v2 tree with the cpu, memory, and pids controllers handed down to \
/sys/fs/cgroup/ffx, and the nftables table that refuses every packet a pool \
uid sends, loopback included, until an allowlist names its destination. The \
rules load from rcS, before the network starts."
LICENSE = "MIT"
LIC_FILES_CHKSUM = "file://${COMMON_LICENSE_DIR}/MIT;md5=0835ade698e0bcf8506ecda2f7b4f302"
SRC_URI = " \
file://forgefirm-sandbox.init \
file://ffx.nft \
"
S = "${WORKDIR}"
inherit update-rc.d useradd
INITSCRIPT_NAME = "forgefirm-sandbox"
# 30 in rcS: /sys is mounted (sysfs.sh, S02) and nothing the script needs
# lives on /data. The network starts in rc5 (S01), so no pool uid ever has
# an interface to send on before the rules are in the kernel.
INITSCRIPT_PARAMS = "start 30 S ."
# The pool. Static ids below 1000, in a block no other account uses (the
# image's dynamic system ids count down from 999): the forgefirm-users render
# replaces only the accounts from 1000 up, so an account reset leaves these
# alone, and the read-only rootfs never needs a useradd at run time. One
# group per account, so no two slots share a group. No home, no shell, and
# the password field useradd leaves locked. The rules in ffx.nft and the
# pool size here name the same range.
FFX_POOL_SIZE = "32"
FFX_POOL_BASE = "800"
USERADD_PACKAGES = "${PN}"
GROUPADD_PARAM:${PN} = "${@'; '.join('--system -g %d ffx%d' % (int(d.getVar('FFX_POOL_BASE')) + n, n) for n in range(int(d.getVar('FFX_POOL_SIZE'))))}"
USERADD_PARAM:${PN} = "${@'; '.join('--system -u %d -g ffx%d -M -d /nonexistent -s /bin/false ffx%d' % (int(d.getVar('FFX_POOL_BASE')) + n, n, n) for n in range(int(d.getVar('FFX_POOL_SIZE'))))}"
do_install() {
install -Dm 0755 ${WORKDIR}/forgefirm-sandbox.init ${D}${sysconfdir}/init.d/forgefirm-sandbox
install -Dm 0644 ${WORKDIR}/ffx.nft ${D}${sysconfdir}/forgefirm/ffx.nft
}
RDEPENDS:${PN} = "nftables"
@@ -64,6 +64,13 @@ IMAGE_INSTALL:append = " grblhal-glowforge forgectrl gfhome gfcloud v4l-utils fw
# forgefirm-persist: the boot timestamp and the random seed on /data.
IMAGE_INSTALL:append = " forgefirm-users forgefirm-hostname forgefirm-banner forgefirm-persist"
# forgefirm-sandbox: what an extension package is held by, in place before
# any package exists: the ffx account pool, the cgroup v2 tree with the cpu,
# memory, and pids controllers, and the nftables rules (nft comes with it)
# that refuse everything a pool uid sends, loaded from rcS before the
# network starts.
IMAGE_INSTALL:append = " forgefirm-sandbox"
# The rootfs mounts read-only on both images; /data (p3) is the writable
# partition. read-only-rootfs is poky's feature for it: the root line of
# /etc/fstab (the BSP's, already ro) and ROOTFS_READ_ONLY in