Commit Graph
221 Commits
Author SHA1 Message Date
ScottW514 1e855ab212 The density floor closes the low end: a commanded 1 percent now marks
$35 = 10 under the density model is a density floor, not a duty floor. It
maps S onto 9.4-100 percent density, so a commanded 1 percent lands at
10.2 percent, just above the marking floor the earlier ladders measured.
A ladder reweighted to the bottom of the user scale - 1, 2, 5, 10, 20,
40, 70, 100 percent of S - marked on all eight rungs, with eight current
segments over a 42.0 s window against exactly 8 x 5.25, and segment means
climbing 136 to 968.

That meets the goal the ladders started from: a user's 1 percent is a
real visible mark rather than silence, and 100 percent is full power. It
took all three pieces - density so every level is real pulses, the
minimum pulse so they stay strikeable, the floor so the user's range sits
on the band that works.

dladder no longer tells the operator to re-run at other base periods to
choose one. The period cancels out of the low end, and what a failing
rung now indicates is a floor set too low.
2026-08-17 21:58:05 -04:00
ScottW514 d7c23cce19 A longer minimum pulse is worse: the gap is what decides striking
min_ticks 6 broke 5 percent striking - seven current segments where 3 gave
eight, on a 36.5 s fire window against 41.4 s for eight rungs. Below the
minimum the model emits min ticks every min/on periods, so the interval
between pulses is min_ticks x tick / density and the base period cancels,
which is also why periods 10, 20 and 40 gave identical results earlier.
At 5 percent that is 2.26 ms at min 3, which struck, against 4.51 ms at 6,
which did not: doubling the minimum doubles the gap as well as the pulse,
and the discharge is re-struck each pulse.

min 3 sits at the factory's own operating point - its 6.5 percent engrave
jobs place 100 us pulses 1.54 ms apart against 1.64 ms for min 3 at that
density - and 6 is outside anything the factory does. The bench is back
at 3.

Measured band for this tube: strikes from ~5 percent density, marks from
~10 percent at F300. That closes the pulse-structure route to a usable
1 percent, since the interval grows as 1/density and 1 percent implies an
11 ms gap. The low end is a scaling problem, and $35 is the control.
2026-08-17 21:52:15 -04:00
ScottW514 ff454ed537 Retract the pulse-length conclusion; show the minimum in the drill table
The fifth ladder, the first with a minimum pulse, moved the floor down a
full rung: only 5 percent failed to mark, and 5 percent now strikes. The
trace carries eight current segments where the run before it had seven,
with fire beginning at 6.2 s exactly at rung 1 and the usual flat
saturated final segment anchoring the count from the other end.

That retracts what the previous entry concluded. Pulse length is not
irrelevant: 10 percent moved from no mark at F100, with three times the
dose per millimeter, to a mark at F300 at the same density, the only
change being its pulses growing from 36-71 us stubs to 106 us. The
matched-pairs argument was sound but drawn entirely from comparisons at
or above 20 percent density, where every pulse length in play was already
long enough - it generalized from the one regime where pulse length does
not bite. Above ~100 us dose governs; below it pulse length does; below
~36 us the supply does not strike. The factory's 100 us quantum sits on
that boundary.

dladder now reads laser_pulse_min_ticks and prints what is actually
emitted. Without that its table reports the pulse density alone would
give, which is wrong wherever the minimum applies - at min 6 the bottom
four rungs all emit 213 us and vary their rate instead, and the operator
reads that table to interpret the material.
2026-08-17 21:47:40 -04:00
ScottW514 a3e83c4fd5 Cover the minimum pulse width; record what the density ladders measured
Rule 15 in the stream harness holds both halves of the minimum
(grblHAL-glowforge f7e8c17, pinned here): no emitted burst falls below
laser_pulse_min_ticks, excepting one clipped by fire going off mid-burst,
and the levels too faint to fill a window still render their exact
average density. Checked against a run at minimum 1 so it cannot pass
vacuously - level 2 goes from 444 bursts of one tick to 147 of three at
the same density, and levels already above the minimum are unchanged.

Four bench ladders settle the base period at 20. The same six rungs
marked in all of them, and the matched pairs across periods separate the
variables: at identical pulse length, halving the density killed the
mark; at identical density, varying the pulse 3x changed nothing. Feed
does not move it either - 10 percent at F100 carries 44 percent more
energy per mm than 20 percent at F300, which marks, and still left
nothing. The low-end marking limit is average power, not dose per length
and not pulse length.

The F100 trace separates two failures that look alike on the material:
seven current segments for eight rungs, anchored by a flat saturated
final segment that can only be full density, put 5 percent at no
discharge at all and 10 percent at a full 15 seconds of current with no
mark. Only the first is ours, and the minimum pulse is the answer to it.

Also recorded: the factory's Precision Power 1 runs a 19.53 percent FIRE
duty cycle at full PWM duty, and its 1-100 scale maps onto density
18.9-79.5 percent, so its 1 percent is the bottom of the useful band
rather than 1 percent of the range. Under the density model $35 and $36
are that same control - a density floor and ceiling - which makes the
user-facing scale a settings choice rather than new code.
2026-08-17 21:37:41 -04:00
ScottW514 c2c627d9b6 Prove the density dose model host-side; close the idle-gap level loss
Four new sessions in the stream harness cover the model (grblHAL-glowforge
2bca017, pinned here). Density renders the commanded level exactly -
levels 2, 3, 7, 15, 25 and 38 came back as 0.0158, 0.0237, 0.0551,
0.1182, 0.1969 and 0.2993 against level/127 of 0.01575, 0.02362, 0.05512,
0.11811, 0.19685 and 0.29921 - S1000 renders 1.0000 and still ends dark,
every power byte carries full duty, and a level change inside a run costs
no stream byte where analog ships one per level.

Rule 13 is the one worth having: the same job run under both models
produces an identical motion grid tick for tick, and all 20051 density
FIRE ticks fall inside the 169776 the analog run fired. The model masks
the core's fire state and never sources one, measured rather than argued.

Rule 14 covers the idle-gap fix: a standalone S between moves, from a
sender slow enough to drain the planner, now fires each move at its own
level (28338 ticks each at duties 30, 52 and 84). Before the fix duty 30
held all 85014 and the other two levels never appeared. That closes
"Next work" item 18, which this work opened earlier today.

The harness now derives its expectations from the board's floor and
chains two launches over one settings file, because the core precomputes
the S to duty mapping once when the spindle is enabled: $35 written at
runtime persists and reports immediately but only enters force at the
next controller start. That is recorded in BRINGUP beside the existing
defaults note, and laser.power-floor's failure message now says so.

Acceptance: the density path shipping off by default is inert until
laser_power_model is set, and the laser tests' covers already name
grblhal-glowforge src/**; the model's own acceptance test waits on the
bench drill that picks the base period.
2026-08-17 20:24:43 -04:00
ScottW514 cb41a6030c Commission the laser duty floor; record how the factory sets power
The pthresh ladder on scrap puts the tube's two thresholds far apart: the
discharge strikes between 2 and 3 percent duty, but nothing lases usefully
below 16 percent (PWMSAR 20), and the rungs between show only a spot at
each line start. $35 ships at 16 (grblhal-glowforge 9466f76, pinned here).

The drill said current lift-off and first mark share a rung; this run
falsifies that, so its docstring and read-the-material text now name both
thresholds and warn that a start-of-line spot is below the threshold, not
at it.

A start-of-line spot is also what a full-power leak at a kernel run start
would look like, so laser_stream_test gains a ladder session (rule 10):
every FIRE tick must ride a commanded duty, and the fire ticks must divide
evenly across rungs. Both hold exactly - six commanded duties, no others,
and 28296 fire ticks on every rung - so the spots are the tube, not the
stream. The harness now derives its expectations from the floor, which
moves the M4 session's S500 plateau from 63 to 73.

laser.power-floor is a new auto acceptance test, the suite's only
non-firing one: a machine must actually carry the commissioned floor,
since stored settings beat freshly baked defaults.

Three cloud cuts of one square at Precision Power 1, 100 and Full Power
show what the analog path is competing with: the power byte is pinned at
127 in all three, dose is FIRE-bit density on a fixed 7-tick period with
the on-count dithered between adjacent integers, and the power setting
never reaches the machine at all. Facts bank and item 17 carry the
numbers; CAMPAIGN-LOG carries both sessions.
2026-08-17 19:45:19 -04:00
ScottW514 76d43686d0 Record the step-timing campaign
The dated record for the jerky-at-2000-mm/min thread: what was wrong,
what changed, and what the bench measured.

Keeps the two results that matter. The A/B that isolated the cause -
two acceptance runs 90 s apart on one image, camera the only variable,
PASS with a nice-5 CPU hog and FAIL with the camera streaming - and the
LB-GF-OG-FM job re-run on the fix with the video live, clean at
min margin 3.1 ms of 10.

Also records that max behind is not the instrument: it reads 0.0 ms
through a passing run whose real margin fell to 4.9 ms, because the
producer never falls behind its own wakeup epoch. Only the measured
min margin shows the condition.
2026-08-17 18:20:40 -04:00
ScottW514 64a720bbab Pin the producer-lead change; record what the bench measured
Bump the grblHAL pin to the producer lead and margin instrumentation.

Rewrite BRINGUP item 16 to the present state. The scheduling half is
done - the producer runs SCHED_FIFO below the shipper, core_mx carries
priority inheritance, and clamping is reported per run - and two bench
runs 90 s apart on one image show what that does and does not cover: a
nice-5 CPU hog passes clean while the camera streaming clamps 7 runs,
because per-frame cache maintenance over a 4.8 MB non-coherent capture
buffer is kernel-context work no userspace priority can preempt.

Record the margin finding, since it explains why a 4 ms stall was
enough: the queue depth cancels between the shipper's due index and the
producer's base, so the pacing lead is the only slack. Record the
cycle-churn ceiling that caps that lead at 10 ms, and that the re-base
fix is what unlocks more.

Camera gating stays owed, with capture resolution named as the lever
that shortens the stall rather than merely spacing stalls out.
2026-08-17 17:59:11 -04:00
ScottW514 faaa6cb40b Cover step timing under CPU contention; record the laser power model
Bump the grblHAL pin to the real-time producer change.

Add motion.step-timing-under-load: the catalog had nothing that
exercised step generation while userspace competed for the single core,
which is exactly the gap that let the condition go unnoticed - the ring
never runs dry, so cnc/underruns reads 0 through it. The test asserts
the producer and the shipper both hold SCHED_FIFO, then drives
2000 mm/min round trips against a deliberate nice-5 CPU hog and requires
the controller to report no clamped events.

Add the pthresh live-fire drill: a constant-power ladder from 2 % to
30 % of full on scrap. Because $35 is a percent of full duty and the
rungs are percents of $30 with $31 = 0, the lowest rung that marks reads
directly as the $35 value. It needs $35 = 0 for the run, or the floor
lifts every rung and hides the threshold.

Record both open items in BRINGUP. The laser one carries the finding
that the factory never uses duty as a power control - all five firing
jobs in the captured pulse files pin the power byte at 127 and modulate
dose by dithering the FIRE bit at 6.5-18.8 % density - so the captures
cannot supply a $35 default, and the duty to optical-power transfer
function of this supply has never been measured.
2026-08-17 16:45:37 -04:00
ScottW514 c425822b82 docs, acceptance: the cameras as users meet them, and 8 MP at full resolution
docs/VIDEO.md is the user-facing camera guide: what the endpoints return, what
the sensors can do that ForgeFIRM does not send and why, the privacy gate, and
the state of 8 MP support.

The 8 MP (OV8856) capture path now reaches the sensor's full 3264x2448 frame.
Its stock RAW10 full-resolution mode runs the link at 1.44 Gbps/lane and the
i.MX6 CSI-2 D-PHY stops at 1 Gbps; the BSP adds a RAW8 mode that carries the
same frame at half the rate, so an HD machine is no longer limited to the
binned quarter-pixel mode. kas/README.md carries the reasoning, the register
deltas and the factory configuration to fall back to if the receiver will not
lock at 720 Mbps/lane.

Acceptance: camera.sensor-profile expects the new OV8856 geometry, the camera
covers map names src/camhealth.* (the src/cam.* glob does not match it, so the
lint would have gone quiet on an uncovered file at the next pin bump), and a
new auto test camera.frame-health asserts a burst of captures with no frames
the capture queue flagged errored - a real signal about the camera ribbon even
though those frames never reach a client.
2026-08-17 14:34:36 -04:00
ScottW514 d3bab940b3 docs: user-facing guides to motion, laser drive and cooling
Two pages aimed at someone who owns the machine rather than works on it,
written for the documentation site. Nothing here is new behavior - it is
the behavior the machine already has, explained where an owner can find
it instead of spread across a kernel contract, a services contract and
three driver headers.

MOTION.md follows one thread: everything physical comes out of a single
fixed-tick byte stream, so the page starts there - the byte layout, speed
as step density rather than clock, the ring and the two ways to fill it,
and the hardware's own stop, halt and resume-with-waypoint. The laser is
presented as part of that stream rather than beside it, which is what
makes the three contract rules (power before fire, no consecutive power
bytes, end dark) and the persisting duty legible instead of arbitrary.
Then geometry and limits, device ownership and the liveness check the
operator sees, and the two modes in full: GRBL from connection through
the arming sequence, the stop/pause/fault table and homing-or-not; cloud
from the preloaded pulse file through the pause backtrack, the park that
ignores the lid, and the ring's cap on job length. A comparison table and
the motion-related settings close it.

COOLING.md explains the engine as what it is - one owner of the thermal
hardware answering a single question for whichever controller runs - and
gives the reasons behind the numbers rather than just the numbers: why
the flow check heats and measures the downstream rise, why 40 percent is
the duty (below it, convection mimics flow), why the settle gate uses a
split-half mean instead of peak-to-peak, and why one bad reading is a
suspicion rather than a fault. Over-temperature, the two diagnostics
tools and when to run them, the settings, and a situation-to-response
table. The fire watch is described honestly: the lid IR channels are
first of all a photometer for the lid lamp, the gate ships watch-only,
and it is not a fire alarm.

Both pages state what is not implemented - low-temperature gates, TEC
control, a fire watch that acts, limit-switch homing - so nobody plans
around them. Constants come from the sources that own them (the feeder
contract, cool.h, the board header, the services contract), not from
prose. README links both.

Documentation only, no behavior change and no catalog consequence: docs/
is outside every layer and .md is excluded from the layer content hash.
2026-08-17 11:52:41 -04:00
ScottW514 44393f3c11 docs: README updates, remove unused assets 2026-08-17 11:25:05 -04:00
ScottW514 05d68ba0a8 docs: BRINGUP describes the present, CAMPAIGN-LOG carries the dated record
BRINGUP had grown to 3,122 lines in which the same subject was answered
differently depending on where the reader stopped: GATE A "stays open, no
live-fire" in the phase text and closed in the campaign record, the catalog
at 24 tests in one section and 35 in another, several "bench validation
pending" headings over bodies that recorded the pass.

Split by kind rather than by age. BRINGUP (907 lines) is the present state
only - status, bench runbook, laser, lid/interlock/button policy, homing,
forgectrl, diagnostics, logging, release acceptance, the measured facts
bank, and a Next work list of the 15 items that are actually open.
CAMPAIGN-LOG (2,657 lines) takes the dated blocks verbatim, in
chronological order, and is append-only: a correction is a later entry, not
an edit. Its two reading rules are stated up front, since moved text keeps
its original "above"/"below" and its pre-split item numbers.

Facts corrected against the tree while rewriting: the liveness probe gates
at p2p 800, not 500, with the real wedge, noise and jolt figures; the panel
has seven tabs including Logs and is built from src/ui/, not ui.c; the
devserver replaced tools/mock.py; /status reports real head presence; the
estop_halts_motion opt-in is gone; core PR #999 is merged, leaving only the
step_us_min commit fork-only; the update system stands at Phase 5, the
uSDHC pads and the 2026-08-08 kernel batch have shipped, and 8 MP camera
capture is tracked here for the first time.

SAFETY and UPDATE-SYSTEM point their drill records at the log. ACCEPTANCE
records that inheritance is local: there is no import of a published
artifact, so a second bench starts from a full campaign - one of the four
open items the retired tool plan held, the rest of which are now in
BRINGUP's acceptance item.

Documentation only, no behavior change and no catalog consequence: docs/
is outside every layer and .md is excluded from the layer content hash.
2026-08-17 11:19:52 -04:00
ScottW514 1c56427cae SAFETY, LIGHTBURN: what the pause and the cancel mean, now that the button drives them
SAFETY: the chain de-energizes itself behind a pause without being asked -
the charge-pump feed ends with the run, so HV_ENABLE drops with the watchdog
about half a second in, and on the resume it is back ~216 ms before the first
step (pad measurements). A pause is deliberately not a cancel: the latch stays
unlocked and the window open, which is what lets the next press resume the
job, while emission still ends because the stream stops driving FIRE. The
disarm grace counts down through a hold, so a job left paused disarms itself,
and a lid or interlock open while paused takes the cancel path - nothing
resumes past an enclosure opening.

LIGHTBURN: the two things an operator needs before pausing a cut - the resume
restarts from where the deceleration ended and accelerates from a standstill,
which M3 shows as a deeper spot and M4 mostly hides, and a job left paused
disarms itself and asks for the button again. Stop leaves the head where it
stopped; returning to the job start belongs to the lid and interlock policy
alone.
2026-08-17 10:08:03 -04:00
ScottW514 74adaf40b8 BRINGUP: lid / button / interlock parity is done; items 4 and 12 close with it
Item 16 becomes the record of what the machine does rather than a list of
what is left: both modes cancel a job on a lid or interlock open - running
or paused - and return the head to where the job started with the lid still
open; the park ignores the lid; the arm wait cancels with the reason named;
the button pauses and resumes, with the factory's backtrack and lead in cloud
mode and feed hold / cycle start in GRBL; lid_policy = hold keeps the stock
door behavior for senders that want it. Item 4's mid-job Door hold is that
policy now, and item 12's LightBurn door handling is closed by the cancel -
LightBurn never lives in Door on the default path.

The GRBL resume dwell that was left open is decided against, on measurement
rather than on the mark: the chain re-arms within ~3 ms of the resume while
motion restarts ~219 ms later, so there is nothing to cover.

The two planning files at the tree root are merged and removed, as the audit
plans were. What was durable in them lives here now: the factory's own
reaction timings on 2.6.0-2228, the pause/resume behavior of the safety chain
at pad resolution, and why the hardware button latch is what makes the armed
window honest. The factory session log they were written from is archived
under _RESOURCES/.

Item 15 gains the catalog's current shape: 35 tests after the parity work and
the sweep that merged the tests sharing a setup, with the auto tests left
separate.
2026-08-17 10:02:25 -04:00
ScottW514 86ce0419e5 forgetest: cloud tests stay in cloud mode; the page can ignore prerequisites
The cloud job tests (lid-abort, lid-during-button-wait, hunt-lid-open,
pause-resume) run in cloud mode and leave the machine there: enter_cloud
reuses a live session (pid-scoped from the client's own websocket state
lines) and switches once from GRBL mode, declaring the change to the
baseline; nothing switches back. Each test judges the log from its own
window, the print by its own "print [id]: finished" line, and waits the
service's follow-up moves out (wait_quiet) before it ends. hunt-lid-open
restarts the cloud client through the supervisor's stop/start lever for
a fresh connect. The former switch-back is what failed the last bench run
of hunt-lid-open (409 machine is not idle: the service was still
re-finding the head after the lid closed).

The baseline is mode-aware: in cloud mode the client owns the GRBL
controller's init values, the lid lamp, and the position counters; the
mode itself is preserved unless the run declared the change
(Context.mode_changed); controller_mode is never handed back as a bare
setting (a bare write left the persisted mode out of step with the live
one). The undeclared-change restore now uses the captured state, which
the post pass never saw before.

The acceptance page gets an "Ignore prerequisites" switch (remembered by
the browser): POST /start {ignore_requires} starts a test whose requires
are unmet, and the run records the unmet prerequisites in its evidence
and log; they stay required for the release. The cloud tests' requires
no longer chain through cloud.mode-switch.

Proof: tests/test_cloud_suite.py replays the four tests on the bench's
own gfcloud excerpts (fixtures/) and the run loop's emitted pause/resume
lines under the real runner Context against a fake forgectrl; baseline
mode tests and the server override test; the whole suite (98) and the
coverage lint pass. Catalog consequence: cloud.* fingerprints move with
the module; the catalog hash moves with the requires.
2026-08-17 06:42:48 -04:00
ScottW514 9878c8d3b9 forgetest: kernel counters prove the return, and ring residue is a leftover that blocks the jog
The lid-cancel tests (motion.lid-cancel-home, laser.lid-cancel-mid-fire)
now check that the KERNEL position counters returned to the job start,
not only grbl's MPos - grbl's drift read 0.000 while the head had not
moved. The baseline reports unplayed bytes queued in the kernel ring
(cnc/position total minus processed) as a leftover and refuses its
return jog while any exist: a run started on top of them replays them
first, which is what put the head into the rail. BRINGUP item 16 records
the bench finding and the stream-engine fix.
2026-08-16 20:47:32 -04:00
ScottW514 b68d790b71 baseline: take the fresh-boot reference after the controller applied its config
/mode reports controller=running at the spawn, so the reference dump
raced grblHAL's init writes and captured the supervisor's motion-probe
values (motor_lock 0, step_freq 10000, y_mode at the module default)
instead of the resting state. boot_reference() now waits for the
controller's markers (step_freq/motor_lock/y_mode at their fixed values,
bounded 20 s, GRBL mode only) plus a 1 s settle before dumping, retakes
a saved reference that shows the pre-config state while the boot is
still fresh, and marks it otherwise. Tests cover the wait, its timeout,
the non-GRBL no-op, the pre-config recognition and the retake.
2026-08-16 20:00:08 -04:00
ScottW514 914c72f8e5 BRINGUP item 16: the arm-wait cancel is a clean soft reset 2026-08-16 19:31:53 -04:00
ScottW514 6b35f9c534 Arm-wait lid cancel expectations: reset banner and no alarm (harness, laser.arm-wait-lid, LIGHTBURN.md) 2026-08-16 19:31:42 -04:00
ScottW514 c4ea96127c Lid/button parity: harness cases, acceptance tests, operator and safety docs, BRINGUP item 16
laser_lifecycle_test.py: the button toggle (press = Hold, press = Run;
the arming press is not a pause), lid and interlock cancel mid-job with
the return to the job start and no alarm, and lid_policy=hold.

Acceptance catalog: motion.button-hold-resume, motion.lid-cancel-home
(operator: travel job, lid open -> cancel message, banner, no alarm,
autonomous return, Idle at the start), laser.lid-cancel-mid-fire (live:
emission stops in hardware, cancelled, returned, armed=false, latch
locked, hardware button latch SET). covers name the switch and laser
sources explicitly.

LIGHTBURN.md: what the lid, Stop and the button now do; SAFETY.md: the
door policy and the button as a software layer; BRINGUP.md: item 16
records the whole parity change (host-proven, bench validation pending)
and items 4/12 point at it.
2026-08-16 19:26:11 -04:00
ScottW514 36b4a3a5f4 Lid open during the arm wait: harness cases, laser.arm-wait-lid, operator doc
laser_lifecycle_test drives the controller's button wait through the
file-backed switch source (GF_SWITCH_FILE): a press with the lid closed
arms and nothing arms before it; the lid or the interlock loop opening
during the wait cancels the job (reason reported, alarm 3, never armed).

Acceptance catalog: laser.arm-wait-lid (operator kind - no press is
given, so nothing can fire) starts a laser job, has the operator open
the lid at the white-button prompt, and checks the cancel message,
alarm 3, armed=false, the kernel latch locked, no emission, and Idle
after $X. covers names glowforge_laser.c, glowforge_switches.c and
glowforge_switch_map.h explicitly.

LIGHTBURN.md: opening the lid or the interlock loop while the button is
lit cancels the job; a press with the lid open never arms.
2026-08-16 18:36:24 -04:00
ScottW514 cd8a01a3d9 bench: every board-runnable tool ported to the bench page
The remaining bench diagnostics run from forgetest's #bench tab. The
tools that also run from a LAN host share scripts/bench/gfbench.py:
GF_HOST names a remote machine (host mode, sysfs through ssh, Grbl and
forgectrl over the LAN); unset, the tool runs on the board itself
(local mode, sysfs directly, everything on 127.0.0.1), which is how the
page runs them - with GF_HOST=127.0.0.1, the panel token in GF_TOKEN
and their data files under <data>/bench/ (FORGETEST_BENCH_DATA). The
helper also reads a machine setting from forgectrl, or from the settings
file on the board while forgectrl is stopped.

Ported: pwm_sweep / pwm_hold (scope = a takeover; the latch relocked,
the write refused if FIRE or LASER_ON reads active), pwm_stream_test
(PASS/FAIL exit), flow_characterize, flow_recheck_char,
flow_warm_validate and flow_matrix (takeovers: forgectrl owns the
thermal hardware, so the page's takeover replaces the tools' own
controller stop/restart, whose command line predated the supervisor;
results and logs in the bench data directory), flow_sustained,
fan_test, temp_calibrate (dry; watch bounded in seconds; the threshold
and the coolant conversion from the shared code), flow_escalate_drill
(cool_confirm_max_s shortened through forgectrl's settings for the
drill and restored; the setting's minimum is the default budget), and
live_fire_drills (<drill> [S] [F], all six drills, host from GF_HOST,
token from the board). flow_matrix joins the registry. What stays
unported cannot run against the machine at all: the two null-sink CI
harnesses and the .puls decoder.

Runner: a scope tool runs inside the takeover wrapper; the bench
environment above is passed to every tool. Tests: test_bench_registry
(registry <-> scripts/bench consistency, every ported tool builds its
command line, every script compiles, gfbench host/local modes) and the
server test (scope tool takeover, the environment reaching the tool).
Local mode smoke-run on the bench (temp_calibrate watch, setting, token)
from /tmp, removed after.

No catalog consequence: bench tools are not image components (dev-only
forgetest); the acceptance catalog is unchanged.
2026-08-16 16:28:39 -04:00
ScottW514 b51e695fb1 manifest: component pins are not layer content
A component pin bump counted as a platform change: the layer content hash
in the platform identity covered the recipe carrying the SRCREV, the
platform is folded into every acceptance fingerprint, so every image
that carried any component update invalidated the whole catalog (dev
image 20260816191951: every test domain-changed after a one-line
forgectrl bump; the two manifests differ only in
platform.layers.meta-forgefirm). The component entry already identifies
the pinned source file by file; the pin double-counted it.

Component pins now live in <recipe>-pin.inc (SRCREV and the PV that
moves with it, nothing else) - forgectrl, grblhal-glowforge and
forgefirm-app here, the BSP components in meta-openglow - and
forgefirm-image-manifest.bbclass leaves *-pin.inc out of the layer
content (FORGEFIRM_MANIFEST_PIN_SUFFIX). Recipe bodies, patches, config
fragments, init scripts and third-party pins with no manifest entry stay
layer content; a pin written into a recipe body still hashes (the safe
direction). manifest-from-tree.py mirrors the rule and reads pins
through the recipe's requires; test_tree_manifest.py proves both
(pin bump: hash unchanged; recipe body or inline pin: changed).
Bitbake resolves the same SRCREV/PV for every pinned recipe.

Docs: ACCEPTANCE.md (what layer content is), kas/README.md (the pin
files in the push order), BRINGUP.md (the finding and the bench
consequence: the first image built with the pin files is itself a
platform change, so its campaign is a full one; pin bumps inherit
after it).

No catalog consequence: nothing in the image's behavior changes; the
change is to the acceptance identity computation, proven by the unit
tests and the CI lint on the tree manifest.
2026-08-16 16:06:58 -04:00
ScottW514 bf066d4e27 BRINGUP: images 20260816191838 / 20260816191951 built with forgectrl c8f6558 and the day's forgetest 2026-08-16 15:21:55 -04:00
ScottW514 57465d2c00 BRINGUP: item 15 - the acceptance tool is bench-validated; the confirmation run rides the next image 2026-08-16 15:12:48 -04:00
ScottW514 df4581134e BRINGUP: campaign 26/26 (export exercised, not a release); lid lamp setting and probe hardening landed and proven 2026-08-16 15:07:58 -04:00
ScottW514 d0291ec7cb forgetest: the lid lamp idles at forgectrl's lid_lamp_idle; the liveness test masks and restarts
The lid lamp now has a resting policy in forgectrl (lid_lamp_idle,
default 236, asserted at start and at every spawn), so the baseline
expects it there instead of preserving whatever level a boot left, and
forgectrl.settings-bounds proves it: resting at the setting, 256 / -1 /
'bright' refused, a new level applied to the lamp at once, the cleared
key back to the default. Clearing a key goes through the query-string
form (an empty JSON value reads as no setting).

motion.liveness-probe adds the regression the bench needed: with every
axis masked (cnc/motor_lock=15, as a bench tool may leave it) forgectrl
is restarted and its fresh probe must read MOTION OK on the first try -
the probe unmasks the axes itself - and the controller comes up with the
mask cleared. Bench 2026-08-16: MOTION OK at p2p 2047/1341 against the
800 threshold, no ladder.

The baseline's settle no longer counts 'probe verified, spawn pending' as
settled (the post pass ran between the probe's own writes and the
controller's init writes and mis-flagged motor_lock/step_freq): it waits
for the controller to be running, with a bounded allowance for a respawn
backoff.
2026-08-16 15:07:34 -04:00
ScottW514 8654c36e99 BRINGUP: forgetest bench campaign 2026-08-16 - 22/26, motion and cloud green, the live tests remain 2026-08-16 14:22:32 -04:00
ScottW514 5002d59dd7 ACCEPTANCE: the fresh-boot reference is taken after a power cycle; the head-return rule 2026-08-16 14:22:08 -04:00
ScottW514 0312dec22b BRINGUP: forgetest bench campaign 2026-08-16 - 14/26, the baseline rule, the masked-probe finding 2026-08-16 13:36:59 -04:00
ScottW514 4aaedc8080 forgetest: every run starts from, and leaves, the fresh-boot idle state
A baseline pass brackets every test and bench tool: before the run the
machine is verified against the fresh-boot idle state and anything off it
is restored; after the run - pass, fail, or abort - it is restored again.
Fixed items are the resting values the boot establishes (module defaults,
forgectrl's start-up writes, the GRBL controller's init writes) and
forgectrl's idle picture (controller running with motion verified, no
diagnostic, camera and cooling engines idle); preserved items (lid lamp
level, position counters, settings map, controller mode) are captured
before and handed back after. Deviations are leftovers: in the run pane,
in the result's evidence, and on the page - attributed to the previous run
when found before, to the run itself when found after.

forgetest takes a fresh-boot reference once per boot (within ten minutes
of boot, after the supervisor settles) as the session's resting lid-lamp
level and the check on the fixed values; the values were confirmed
against a fresh boot of the dev image on the bench (step_freq rests at
28160, the controller's default tick, not the probe's 10000).

Takeover runs capture the controller-owned kernel attributes on entry and
write them back before forgectrl restarts: the bench found the kernel
tests leaving motor_lock=15 behind, which masked the supervisor's
liveness probe - no motion by construction, a false driver-wedge verdict,
the rail-off ladder, and finally motion-fault. The takeover wrapper also
waits for the supervisor to settle on both sides (moved into baseline).

Catalog consequence: none beyond the runner; the tests' own drills are
unchanged.
2026-08-16 13:35:19 -04:00
ScottW514 e19e7c5304 BRINGUP: images 20260815215236 / 20260815215332 built with the logging catalog and the fast sanitizer 2026-08-15 17:54:47 -04:00
ScottW514 d7d68c556e BRINGUP: the swept legacy logs were deleted from the bench 2026-08-15 17:51:22 -04:00
ScottW514 d17655c3c0 forgetest: logs.routing and logs.level-settings; export gets its own timeout
The routing test proves the path every logger takes and stands for the
emitters and relays in every component (fflog in forgectrl and grblHAL,
the SysLogHandler in the Python apps, the supervisor's per-controller
relay, the daemon's fifo relay, the render step): one logger daemon,
rendered rules and the effective record consistent with /logs, the tree,
the daemon's own emitter line, `logger` probes routed by program name in
the ff_line format, a stray program only in system/, kernel lines, relay
processes, nothing written outside the tree. The level-settings test
covers the validators and the configured-vs-effective / pending_reboot
contract. Both PASS on the bench with the existing logs.tree-tail-export
through the real Runner. That test's export call now uses a 300 s client:
the sanitized export took 13.9 s on the target and tripped the 10 s
default. Pin: forgectrl 4d19e9d (the sanitizer skips passes a line
cannot match; 4x faster). Coverage lint enforced: 0 uncovered.
2026-08-15 17:44:44 -04:00
ScottW514 07141d60fc BRINGUP: remote syslog verified over the real hop (UDP + TCP, collector outage + queued delivery) 2026-08-15 17:24:45 -04:00
ScottW514 2e82b13438 BRINGUP: $H gfhome routing verified; item 14 closed 2026-08-15 16:31:31 -04:00
ScottW514 7baa6a0352 BRINGUP: unified logging bench validation complete (levels, remote, rotation, export, RT, respawn, cloud routing) 2026-08-15 16:26:58 -04:00
ScottW514 16deb4896f forgetest: /fuse-identity is two-factor; the auth test asserts both refusals
forgectrl reveals the fuse identity only to the token AND the physical
button held, so the token alone must answer 403 with the button message.
The test now asserts the no-token refusal and the token-without-button
refusal and never fetches the identity itself (a 200 would have carried
the fuse password into the result log). BRINGUP: the bench campaign on
the flashed dev image, 7 of 24 passed so far.
2026-08-15 16:07:12 -04:00
ScottW514 5e04f4f4cd BRINGUP: forgetest CI is green on the pushed tree 2026-08-15 15:58:56 -04:00
ScottW514 1179d5e7c1 Release acceptance gate: release.sh refuses to sign without a matching artifact
scripts/acceptance-gate.py recomputes every catalog test's domain fingerprint
from /etc/forgefirm-manifest.json inside the release rootfs and requires the
committed releases/v<version>/acceptance.json to carry a matching PASS
(inherited results not core and newer than the invalidate epoch; the artifact
self-hashed; the catalog identical to the tree). release.sh runs it after the
build and stages the artifact as a release asset; FORGEFIRM_ACCEPTANCE_SKIP=1
bypasses loudly. scripts/manifest-from-tree.py builds the same manifest from
the recipe pins with git for CI and the workstation; forgetest-ci.yml runs the
unit tests and enforces the coverage lint (every manifest path covered by some
test). docs/ACCEPTANCE.md is the contract; the coverage currency rule and the
status live in BRINGUP.
2026-08-15 15:57:12 -04:00
ScottW514 da3725c6de BRINGUP: drop the acceptance-tool text that was swept in from the shared working tree
It belongs to work in progress in another session; committing it here was a mistake. The working tree still carries it uncommitted.
2026-08-15 15:28:39 -04:00
ScottW514 304c4e7186 BRINGUP: unified logging bench checks on the flashed dev image 2026-08-15 15:27:48 -04:00
ScottW514 6050c0e703 Unified logging: rsyslog as the system logger, the ForgeFIRM log tree
rsyslog replaces busybox syslogd/klogd (VIRTUAL-RUNTIME_base-utils-syslog,
trimmed PACKAGECONFIG) and becomes the only log writer: the appended
/etc/rsyslog.conf sets the inputs and the ff_line format and includes
the per-logger rules that `forgectrl --render-syslog` renders from the
machine settings at boot. forgefirm-logrotate becomes forgefirm-logging:
render before rsyslog starts (S19), sweep the pre-syslog log files into
/data/forgefirm/legacy-logs once, and rotate the tree at boot and hourly
by rename + HUP instead of copytruncate. Pins bumped to the pushed
forgectrl (syslog emitter, Logs tab, export), grblHAL-glowforge (syslog
emitter) and python3-gfhardware apps (syslog handlers, capture dir);
the CI harnesses set FFLOG_STDERR=1 so failure diagnostics keep the
controller's log lines. BRINGUP carries the bench validation checklist
(Next work item 14); this is an image change and rides the next flash.
2026-08-15 14:37:33 -04:00
ScottW514 b08e5ab929 BRINGUP: clarify the pre-rename estop references in dated records 2026-08-15 14:16:34 -04:00
ScottW514 3f81597c93 SAFETY: hv_enable run-boundary behavior is established; drop it from the open items 2026-08-15 14:11:32 -04:00
ScottW514 7319c718fe SAFETY/BRINGUP: HV watchdog one-shot measured pulse-to-drop, t_w = 454 +/- 3 ms 2026-08-15 14:01:46 -04:00
ScottW514 d7cb506c46 BRINGUP: hv_enable readback bench-validated on image 20260815162923 2026-08-15 13:36:51 -04:00
ScottW514 87a8ccbd59 BRINGUP: image 20260815162923 built on the hv_enable pins; flash pending 2026-08-15 12:41:52 -04:00
ScottW514 f75629a4df hv_enable: EV_SW bit 4 is the HV_ENABLE readback; pins for the rename
SAFETY.md and the safing figure name GPIO4_06 for what it is - the
readback of the chain's HV_ENABLE output through U24 (the factory net
label E-STOP is kept as a note); BRINGUP records the rename, the
device-tree polarity flip that makes bit 4 read as HV_ENABLE itself,
the removal of the estop_halts_motion opt-in, and the bench check for
the flash that ships it. Recipe pins move to forgectrl 801f1f3,
grblHAL-glowforge b629c18 and python3-gfhardware c3d1790 (PV 0.1.5).
2026-08-15 12:28:58 -04:00