Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.
Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.
scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
- a failure while building the response (a non-finite number included) is a 500 inside the
envelope, never a 422 or a render crash outside it
- an engine ValueError keeps its message but is released and raised unchained, like an OOM
- the prompt is built (and the model's own validation run) before the forward
- the row cap is counted before any ordering is built
- /health reads a device name cached at load, so it makes no driver call off the inference thread
torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the
per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on
fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the
10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
Its dataclasses use postponed annotations and look their module up in sys.modules while the
class is built; importing by file path without registering it failed startup (closed).
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s
9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with
expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the
recreated container's own env. Transcripts: 95.8% word-sequence similarity
vs 300 s, diffs mostly casing/punctuation.
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
Both verdicts now build the n for the U11b data-deletion gate (three
consecutive PASS batches at off). First real run 20260930T090608Z via the
detached-worktree path: PASS, user median 0.83.
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
For infra-hermes's daily batch through the U11 off window (operator ruling
2026-09-29 2340; mode revised to off 2026-09-30). Enforces worldtree-dev's
terms: flock plus a refusal while any harness batch runs, code under test =
demo's deployed sha (a detached worktree when the main tree has moved; refuses
on pyproject/uv.lock/packages drift the shared venv cannot honour), no git
commit, no retry. Exit 0 PASS / 1 FAIL / 2 harness error / 3 refused.
The 2026-09-28 expansion grew the one-page design to nine pages by
bolting a nav band under each hero. This makes the cluster one site:
- One site header on every page: the wordmark (home link) and a single
row of short, consistent labels (Applications, Systems, Hosting, AI,
Engagement, Track record, Contact). The nav no longer wraps into two
rows at desktop width or four on a phone.
- A real <h1> on every page (the title was a <p>), a skip link, and one
footer outside <main> on every page, home included.
- The contact page was unreadable: its section borrowed the home page's
dark contact-card id, so its copy was grey on near-black (2.3:1) and
its headings measured 1.0:1. It is an ordinary section now.
- Section numbers only where there is a sequence (home, engagement).
- The spec bar names all four services; it stacks cleanly on a phone.
- Contrast: --text-tertiary 4.3:1 -> 4.96:1, --accent-text 4.36:1 ->
4.74:1 on paper. Track-record labels drop tracked capitals.
- The grid behind the hero and the contact card is its own layer; the
contact card loses its left side-tab stripe.
- The stylesheet link is versioned (style.css?v=2026-09-28): CSS is
cached for an hour, and the new markup must not meet the old styles.
- README: nine pages, the shared structure, the CSS-version rule, and the
operator's anti-slop waivers (the hero's volt rule, Space Grotesk).
Copy is unchanged apart from nav labels (the brief's legal constraints
hold). Still zero JavaScript and zero external requests. Impeccable
detector (controls 7/0, static + rendered at 1280 and 390, light and
dark): 200 unwaived -> 54, all of them the two waived brand devices;
two identical runs. Operator: "ship the site cleanup, keep the rule and
the font".
Blender is now a mandatory stage in draupnir's pipeline (Prime, 2026-09-28), and draupnir asked
for eight add-ons from extensions.blender.org: SurfacePsycho 0.10.4, CAD Sketcher 0.32.1,
3D-Print Toolbox 1.4.1, STEP Importer 1.2.1, Bool Tool 2.1.0, LoopTools 4.7.7, MeasureIt 1.8.4,
3MF Import/Export 2.7.7.
- stacks/blender/extensions.lock pins each by version and archive sha256.
- scripts/blender-extensions sync builds fv-ml1:/tank/blender-extensions/5.2/system with Blender's
own install-file, pre-warms and byte-compiles it, checks a read-only enable, then swaps it in.
It refuses while the GUI or a blender-run job holds the old directory.
- conf/scripts/startup/fleet_extensions.py enables every package in the System repo: in a timer
in the GUI (after the prefs load), and as --python ahead of the caller's args in
blender-run --extensions (a failed enable exits 1 before the caller's script).
- It also patches SurfacePsycho's sp_overwrite_segment_selection from eval() to literal_eval():
the eval walked past MCP safe mode (control: unpatched ran code, patched refuses).
- blender-run: --extensions (bind mounts via --mount so a missing source fails instead of being
created); USER/LOGNAME set, which CAD Sketcher's getpass needs.
- compose.yaml mounts the repo read-only and the hook into the GUI container. NOT yet deployed.
- scripts/blender-probes/extensions_acceptance.py: one operator run per add-on, safe-mode
compliant. Headless 8/9 online and with --network none; CAD Sketcher sketching is GUI-only.
A Python audit hook saw no network/process events (positive control fired).
Single page grows to nine: custom app development, systems engineering,
hosting, and AI services each get a detail page, plus engagement,
track-record, contact, and privacy. Cross-linked nav and footer on every
page; brand CSS extended in place (still zero external requests, zero JS).
Motivation: an app-store registration review flagged the single-page site
as minimal content. This version is intended to be reverted after approval.
draupnir review: the remote job dir was keyed on the basename alone, so two
local dirs with the same name shared one remote dir, and a rerun inherited
stale files. The dir is now <basename>-<8 hex of sha256(abs path)>, and it is
mirrored with rsync --delete, confined to that one directory. There is no rm on
a computed path. Verified: a file deleted locally is gone from the rerun's
remote dir. The /defaults stderr line is documented as harmless noise.
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).
Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.