althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so
Restart=on-failure never revived it; hermes-gateway sat pull-only for five
days with a yellow verdict nobody was watching. Fixed the death mode with a
Restart=always drop-in (also on the jekyll twin, same latent bug) and added
this watchdog for every other death mode: 5-min timer, alert on the second
consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery
mail on green, exit 2 distinguishes a broken alarm wire from a down seat
(backup-freshness precedent). Lifecycle exercised end-to-end against a dead
port before install.
Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85%
twice in 14 h. The swap partition sat right after sda1 and blocked growth,
so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows
sda1 in place (start sector unchanged) and ext4 online, and sets initramfs
RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%.
The in-browser recorder calls getUserMedia, which browsers refuse on http://
origins, so it sat at "Initializing recorder...". scriberr.nh3.phasefinal.com
is now fronted by the fleet TLS caddy on nh3-dev (wildcard cert) and added to
Scriberr's ALLOWED_ORIGINS; the Homepage link points at it. The plain
http://10.251.50.54:8080 URL keeps working except for recording.
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.
scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.
Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).
The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:
- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
returns to ~2,108 MiB rest after a 12-min file instead of parking at the
peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
answer 503 with the seat alive (proved at cap=2000), so the failure lands
on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
free (illegal-memory-access wedge when both were on first try). Cost:
12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
under the sherpa seat.
Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.
GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.
Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.
Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.
License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).
- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
(paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
-3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.
Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.
Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.
scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).
Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.
Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.
Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.
scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
Scriberr moved to fv-ml1 GPU 3 on 2026-09-30, so the card is no longer idle when the rebuild runs. The memory stage now requires >= 20 GB free instead of an idle card (GPUs 0-2 still fail that) and attributes the peak only to the host PIDs of its own container, captured with docker top alongside the 0.2 s nvidia-smi samples. Verified with five processes on the card: peak 5,496 MiB, identical to the exclusive-card figure.
Audit finding on 0.1.1: a Jev client that always sends an images array with no
images was rejected for nothing. Only a non-empty value is 422 now. Also aligns
the contract's response example with the wire (model is the name@revision
string, not an object).
Re-accepted live: images []/null -> 200, ["a.png"] -> 422; JevBench all 202/231,
hard 83/111, 0 changed rows across the bench's r1..r4 (924).
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).
Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.
Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.
scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
- a failure while building the response (a non-finite number included) is a 500 inside the
envelope, never a 422 or a render crash outside it
- an engine ValueError keeps its message but is released and raised unchained, like an OOM
- the prompt is built (and the model's own validation run) before the forward
- the row cap is counted before any ordering is built
- /health reads a device name cached at load, so it makes no driver call off the inference thread
torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the
per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on
fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the
10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
Its dataclasses use postponed annotations and look their module up in sys.modules while the
class is built; importing by file path without registering it failed startup (closed).
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s
9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with
expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the
recreated container's own env. Transcripts: 95.8% word-sequence similarity
vs 300 s, diffs mostly casing/punctuation.
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
Both verdicts now build the n for the U11b data-deletion gate (three
consecutive PASS batches at off). First real run 20260930T090608Z via the
detached-worktree path: PASS, user median 0.83.
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
For infra-hermes's daily batch through the U11 off window (operator ruling
2026-09-29 2340; mode revised to off 2026-09-30). Enforces worldtree-dev's
terms: flock plus a refusal while any harness batch runs, code under test =
demo's deployed sha (a detached worktree when the main tree has moved; refuses
on pyproject/uv.lock/packages drift the shared venv cannot honour), no git
commit, no retry. Exit 0 PASS / 1 FAIL / 2 harness error / 3 refused.
The 2026-09-28 expansion grew the one-page design to nine pages by
bolting a nav band under each hero. This makes the cluster one site:
- One site header on every page: the wordmark (home link) and a single
row of short, consistent labels (Applications, Systems, Hosting, AI,
Engagement, Track record, Contact). The nav no longer wraps into two
rows at desktop width or four on a phone.
- A real <h1> on every page (the title was a <p>), a skip link, and one
footer outside <main> on every page, home included.
- The contact page was unreadable: its section borrowed the home page's
dark contact-card id, so its copy was grey on near-black (2.3:1) and
its headings measured 1.0:1. It is an ordinary section now.
- Section numbers only where there is a sequence (home, engagement).
- The spec bar names all four services; it stacks cleanly on a phone.
- Contrast: --text-tertiary 4.3:1 -> 4.96:1, --accent-text 4.36:1 ->
4.74:1 on paper. Track-record labels drop tracked capitals.
- The grid behind the hero and the contact card is its own layer; the
contact card loses its left side-tab stripe.
- The stylesheet link is versioned (style.css?v=2026-09-28): CSS is
cached for an hour, and the new markup must not meet the old styles.
- README: nine pages, the shared structure, the CSS-version rule, and the
operator's anti-slop waivers (the hero's volt rule, Space Grotesk).
Copy is unchanged apart from nav labels (the brief's legal constraints
hold). Still zero JavaScript and zero external requests. Impeccable
detector (controls 7/0, static + rendered at 1280 and 390, light and
dark): 200 unwaived -> 54, all of them the two waived brand devices;
two identical runs. Operator: "ship the site cleanup, keep the rule and
the font".
Blender is now a mandatory stage in draupnir's pipeline (Prime, 2026-09-28), and draupnir asked
for eight add-ons from extensions.blender.org: SurfacePsycho 0.10.4, CAD Sketcher 0.32.1,
3D-Print Toolbox 1.4.1, STEP Importer 1.2.1, Bool Tool 2.1.0, LoopTools 4.7.7, MeasureIt 1.8.4,
3MF Import/Export 2.7.7.
- stacks/blender/extensions.lock pins each by version and archive sha256.
- scripts/blender-extensions sync builds fv-ml1:/tank/blender-extensions/5.2/system with Blender's
own install-file, pre-warms and byte-compiles it, checks a read-only enable, then swaps it in.
It refuses while the GUI or a blender-run job holds the old directory.
- conf/scripts/startup/fleet_extensions.py enables every package in the System repo: in a timer
in the GUI (after the prefs load), and as --python ahead of the caller's args in
blender-run --extensions (a failed enable exits 1 before the caller's script).
- It also patches SurfacePsycho's sp_overwrite_segment_selection from eval() to literal_eval():
the eval walked past MCP safe mode (control: unpatched ran code, patched refuses).
- blender-run: --extensions (bind mounts via --mount so a missing source fails instead of being
created); USER/LOGNAME set, which CAD Sketcher's getpass needs.
- compose.yaml mounts the repo read-only and the hook into the GUI container. NOT yet deployed.
- scripts/blender-probes/extensions_acceptance.py: one operator run per add-on, safe-mode
compliant. Headless 8/9 online and with --network none; CAD Sketcher sketching is GUI-only.
A Python audit hook saw no network/process events (positive control fired).
Single page grows to nine: custom app development, systems engineering,
hosting, and AI services each get a detail page, plus engagement,
track-record, contact, and privacy. Cross-linked nav and footer on every
page; brand CSS extended in place (still zero external requests, zero JS).
Motivation: an app-store registration review flagged the single-page site
as minimal content. This version is intended to be reverted after approval.
draupnir review: the remote job dir was keyed on the basename alone, so two
local dirs with the same name shared one remote dir, and a rerun inherited
stale files. The dir is now <basename>-<8 hex of sha256(abs path)>, and it is
mirrored with rsync --delete, confined to that one directory. There is no rm on
a computed path. Verified: a file deleted locally is gone from the rerun's
remote dir. The /defaults stderr line is documented as harmless noise.
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).
Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
The MCP server (mcp-for-blender 2.1.1, frozen requirements) runs inside the Blender
container. Its add-on is vendored at upstream 41a18432 (MIT) and started by a
startup hook. scripts/blender-mcp carries the stdio over ssh + docker exec, so the
add-on socket, which runs arbitrary Python with no auth, stays on the container's
localhost with no published port. It also runs there because viewport screenshots
need a filesystem shared by server and Blender. Telemetry is off and safe mode is
on. The hook also defaults Cycles to OptiX on GPU 3, because safe mode forbids
agents from touching preferences.
Verified end to end from nh3-dev: 36 tools; a GPU render of an agent-built scene;
a viewport screenshot; and safe mode refusing 'import os'. Blender left down
(on demand).
Uses linuxserver/blender (Selkies Wayland desktop, NVENC), digest-pinned. The
container sees only GPU 3, has restart "no", and runs only while in use,
because GPU 3 is the reserve card for a full-size vLLM seat. Web desktop on
:3001 with basic auth (vault fv-ml1/blender-web-password). /work is on /tank
and is not backed up; /config lives under /opt/docker (restic).
Acceptance: Cycles finds the card on OptiX and CUDA (sm_120 kernels ship in the
build). The heavy self-test renders in 3.54 s on OptiX vs 22.71 s on CPU, a
functional check with n=1. Web auth answers 401 without credentials and with a
wrong password, and 200 with the right one.