Commit Graph
100 Commits
Author SHA1 Message Date
vh 36c52d8860 docs(meshcentral): stale CIRA tunnel reproduced on a single shutdown 2026-10-02 22:53:30 -07:00
vh ad0322bf83 ops(meshcentral): diagnose and fix HW Connect stuck at Setup (stale CIRA tunnel); add relay probe 2026-10-02 22:49:27 -07:00
vh 55004090d1 ops(esh-pve-2): register host, AMT phoning home to MeshCentral; note MeshCentral first-CIRA crash race 2026-10-02 22:25:37 -07:00
vh e707d87713 feat(esh-pve-2): drop stale 1 TB boot entry, wipe it into LVM-thin storage vmstore 2026-10-02 22:08:19 -07:00
vh 6315ce38d5 memory: ESH tank verified residual, cleared, full scrub running 2026-10-02 19:26:30 -07:00
vh de143364fd docs: fleet Proxmox host inventory snapshot 2026-10-02 (all sites) 2026-10-02 19:23:58 -07:00
vh 62254023f3 docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding 2026-10-02 19:20:28 -07:00
vh 5439d5fdc1 albok-service: 0.1.1 (health fix) on nh3-docker 2026-10-02 18:04:57 -07:00
vh d3958655c1 feat(albok-service): deploy the fleet knowledgebase service on nh3-docker
albok-service 0.1.0 (vh/albok 0b37431), image pfi/albok-service pinned by
digest, published on host port 8392 because 8390 is the post office.
Host prep playbook creates the fixed ids (albok 1500, albok-read 1510,
albok-personal 1511) and the local store/private roots; the container
gets a mounted /etc/group and group_add so the service can resolve and
chgrp its wing dirs. The config carries a LiteLLM key scoped to
qwen3-embedding and lives outside the deploy-synced conf dir. DNS name
albok.nh3.internal.
2026-10-02 17:56:34 -07:00
vh 9c233ca255 memory: nh3-pve-2 end-of-day state and on-site reinstall checklist 2026-10-02 17:40:50 -07:00
vh 56c372dd6f docs(pfi-tacticalrmm): rmm-mesh nginx upload limit raised to 4G (was the 1 MB default) 2026-10-02 17:08:21 -07:00
vh 5c0d5f0a73 docs(pfi-tacticalrmm): MeshCentral site-admin account lkraven 2026-10-02 17:06:19 -07:00
vh 34659ae9e7 ops(nh3-pve-2): AMT 21 configured (KVM, no-consent, listener) and phoning home to MeshCentral 2026-10-02 14:54:56 -07:00
vh e73adfe7ae dns+memory: nh3-pve-2-amt 10.100.250.63 (UDM port 5 on nh3-mgmt, reservation) 2026-10-02 14:43:16 -07:00
vh 59c9c0715a memory: MS-03 roles — esh-dev inherits nh3-dev's sessions; nh3-pve-2 purpose TBD by design 2026-10-02 09:17:39 -07:00
vh 536230170e memory: one MS-03 deploys as nh3-pve-2 (AMT on DHCP for phone-home; addressing plan) 2026-10-02 09:10:31 -07:00
vh 9dba6baccf homepage: NH3-PVE-AMT card opens MeshCentral (AMT now phones home; LAN management is off) 2026-10-02 09:06:49 -07:00
vh 1a2d763de4 ops(nh3-pve): AMT phones home to MeshCentral (CIRA); AMT on DHCP; LAN management now dark by design 2026-10-02 09:06:24 -07:00
vh 9bcb9417fb feat(amt): amt-cira-setup.py — configure Intel AMT phone-home (CIRA) to MeshCentral over WS-Man
MeshCentral only pushes CIRA through an on-host agent, so this mirrors its
amtmanager.js sequence directly over the LAN: trust MeshCentral's root,
add the MPS (username = 16-char meshid prefix), a periodic policy,
BIOS+OS user-initiated connections and a random environment-detection
domain. Idempotent, with a read-only state report. Enumerate uses
Enumerate+Pull (AMT 16 ignores OptimizeEnumeration; caught by a
positive control that first read 0 instances of a class that has one).

Applied to nh3-pve's AMT 2026-10-02 alongside MeshCentral mpsPass and an
ana-gw VIP/policy for 4433; the AMT does not dial out yet (static IP).
2026-10-02 08:50:47 -07:00
vh 61ba384d28 memory: demo fix-forward 7ab6ae40 verified; deploy-gate gap filed as Worldtree #423 2026-10-02 08:24:33 -07:00
vh 1219cfaca9 memory: demo rolled back from broken c2d87263 to fa8bc51c; b193 retired-key list for demo 2026-10-02 08:17:24 -07:00
vh 8b81579bca ops(pfi-tacticalrmm): MeshCentral to hybrid mode; nh3-pve AMT added and connected 2026-10-02 08:14:19 -07:00
vh a967bb95a6 docs(pfi-tacticalrmm): MeshCentral facts — WAN-only mode silently drops AMT adds; CLI access via the vaulted login token 2026-10-02 08:09:02 -07:00
vh cd30592e55 fix(wt-memory-gate-batch): pass --legacy-mode only when the deployed harness has it
b193 (U11b) retired the legacy plane and removed --legacy-mode from the
harness; passing it there fails the daily batch. The wrapper now greps
core/memory_acceptance at the deployed sha and omits the flag when it is
gone, so it is correct on both sides of the b192 -> b193 deploy (checked:
7a83f2f -> passes it, c2d87263 -> omits it).
2026-10-02 07:59:03 -07:00
vh 1cc8ad3292 memory: U11b step 5 done — live legacy memory deleted on demo and personal (H2 0), stamp sent 2026-10-02 07:54:50 -07:00
vh b0789f4b29 memory: U11b — gate-batch wrapper drops --legacy-mode at b193 2026-10-02 07:50:38 -07:00
vh 6e15f610bc memory: U11b gate streak 3 of 3; step 5 held for Prime's direct go 2026-10-02 07:50:12 -07:00
vh 59fb6b77a8 ops(nh3-dev): seat healthz watchdog — the 5-day pull-only outage must page, not age
althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so
Restart=on-failure never revived it; hermes-gateway sat pull-only for five
days with a yellow verdict nobody was watching. Fixed the death mode with a
Restart=always drop-in (also on the jekyll twin, same latent bug) and added
this watchdog for every other death mode: 5-min timer, alert on the second
consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery
mail on green, exit 2 distinguishes a broken alarm wire from a down seat
(backup-freshness precedent). Lifecycle exercised end-to-end against a dead
port before install.
2026-10-02 07:47:17 -07:00
vh 902e16630f feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)
Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
2026-10-02 07:42:45 -07:00
vh 49c8fdf383 memory: nh3-dev snapshot waits for a natural reboot (Prime); post-boot checks listed 2026-10-02 07:36:15 -07:00
vh 6ec76bf35c ops(nh3-dev): grow root into the 378 GB disk online; swap moved to /swapfile
Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85%
twice in 14 h. The swap partition sat right after sda1 and blocked growth,
so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows
sda1 in place (start sector unchanged) and ext4 online, and sets initramfs
RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%.
2026-10-02 07:33:36 -07:00
vh 01c356e7fb docs(ana-docker): record the RustDesk server key and where the client setup steps live 2026-10-01 15:31:32 -07:00
vh abcf4c3166 fix(scriberr): serve over HTTPS via the fleet TLS caddy so the browser recorder works
The in-browser recorder calls getUserMedia, which browsers refuse on http://
origins, so it sat at "Initializing recorder...". scriberr.nh3.phasefinal.com
is now fronted by the fleet TLS caddy on nh3-dev (wildcard cert) and added to
Scriberr's ALLOWED_ORIGINS; the Homepage link points at it. The plain
http://10.251.50.54:8080 URL keeps working except for recording.
2026-10-01 13:07:18 -07:00
vh 474f78811c memory: U11b — writer/reader default to enabled at b193 (no instance action) 2026-10-01 12:35:57 -07:00
vh f37a50c11a memory: U11b C5 — memory_extractor and memory.escalation cleared to retire after b193 2026-10-01 11:25:01 -07:00
vh aa440708c4 memory: U11b C5 post-b193 config retirements listed; memory_extractor held open (escalation.model_role) 2026-10-01 11:23:32 -07:00
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00
vh e5536784e0 memory: U11b gate verdicts arrive from hermes-gateway; reply to infra-hermes 2026-10-01 07:49:49 -07:00
vh 23cedb3259 memory: Worldtree U11b gate streak 2 of 3 (20261001T143101Z PASS) 2026-10-01 07:49:28 -07:00
vh de11e00dfb fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
2026-10-01 05:59:07 -07:00
vh 9af5a16d07 memory: snapshot — dev-backup retention broken since 07-18 (log rotated, errors counted; Prime to rule); gen-small KV + graphs-off left as is 2026-10-01 05:29:27 -07:00
vh 0d9bb109b8 memory: gen-small OOM fixed by parakeet-nemo 0.1.1 (cache return + hard cap); Gitea webhook allow-list lesson 2026-10-01 04:42:13 -07:00
vh 963f9ed8c0 fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:

- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
  returns to ~2,108 MiB rest after a 12-min file instead of parking at the
  peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
  answer 503 with the seat alive (proved at cap=2000), so the failure lands
  on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
  free (illegal-memory-access wedge when both were on first try). Cost:
  12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
  under the sherpa seat.

Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.
2026-10-01 04:39:39 -07:00
vh a816443acd memory: arbo webhook repointed off the retired wg0 IP 2026-10-01 04:30:03 -07:00
vh f83e35b94b memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived 2026-10-01 04:24:44 -07:00
vh 6b66207b0c scriberr: no upstream contribution (Prime) — drop prepared PR text; patches carried locally 2026-10-01 04:20:17 -07:00
vh 03826d2029 docs: gen-small .env.example records util 0.33 + boot-check rule; parakeet-nemo compose note corrected (slack, not KV; steady state 3,582 MiB) 2026-10-01 01:38:42 -07:00
vh cb28d8c951 memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state 2026-10-01 01:37:49 -07:00
vh 392660bd5f docs(parakeet-nemo): steady-state 3,582 MiB and the boot-check arithmetic rule (audit findings) 2026-10-01 01:37:47 -07:00
vh 41d2014df9 docs(parakeet): mark the sherpa seat SUPERSEDED; keep it as the rollback build source 2026-10-01 01:32:57 -07:00
vh de6ea32f34 feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
2026-10-01 01:32:48 -07:00
vh dbd583d6ca memory: irv-ml1 storetank reclaim closed (Prime via comfy-dev); symlink-target lesson 2026-10-01 00:11:05 -07:00
vh aa7eeff445 memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits 2026-09-30 23:54:54 -07:00
vh 1ae324d576 memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived 2026-09-30 23:52:27 -07:00
vh 212b736836 memory: parakeet seat A/B done — runtime is the bottleneck; unified-en NeMo bf16 wins speed+WER; seat defects 2026-09-30 18:53:04 -07:00
vh b38ec6dea0 docs(parakeet): seat A/B - clock times in Pacific military time 2026-09-30 18:51:59 -07:00
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00
vh 11174ffea1 memory: GPU 0 room can come from trimming gen-small KV (Prime); donor analysis 2026-09-30 16:08:10 -07:00
vh aa17f64bd1 memory: warm-up cache audit passed; parakeet seat A/B in flight 2026-09-30 16:05:56 -07:00
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00
vh f65e27b08f feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
2026-09-30 16:02:16 -07:00
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00
vh 1189adbf18 memory: intern-decision warm-up cache tasked to infra-hermes 2026-09-30 15:44:45 -07:00
vh af450f4ef7 docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up 2026-09-30 15:41:55 -07:00
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
vh 92501a29c1 fix(scriberr-rebuild): memory stage works on a shared GPU 3
Scriberr moved to fv-ml1 GPU 3 on 2026-09-30, so the card is no longer idle when the rebuild runs. The memory stage now requires >= 20 GB free instead of an idle card (GPUs 0-2 still fail that) and attributes the peak only to the host PIDs of its own container, captured with docker top alongside the 0.2 s nvidia-smi samples. Verified with five processes on the card: peak 5,496 MiB, identical to the exclusive-card figure.
2026-09-30 13:30:29 -07:00
vh e3dbb08d84 memory: true Jev is text-only (official docs); our real Jev gap is context (7,168 vs 64k tokens) 2026-09-30 13:09:44 -07:00
vh 1faadb45f6 memory: intern-decision 0.1.2 live, re-audit passed 2026-09-30 13:06:53 -07:00
vh 1866c003e8 fix(intern-decision-serve): 0.1.2 treats empty/null images as absent
Audit finding on 0.1.1: a Jev client that always sends an images array with no
images was rejected for nothing. Only a non-empty value is 422 now. Also aligns
the contract's response example with the wire (model is the name@revision
string, not an object).

Re-accepted live: images []/null -> 200, ["a.png"] -> 422; JevBench all 202/231,
hard 83/111, 0 changed rows across the bench's r1..r4 (924).
2026-09-30 13:04:31 -07:00
vh da50696fe8 memory: intern-decision /v1/systemone live (0.1.1), infra-ops audit passed; 2 low findings to infra-hermes 2026-09-30 13:01:05 -07:00
vh ff552abf1f feat(intern-decision-serve): 0.1.1 adds POST /v1/systemone (Jev wire shape)
Straight passthrough to the checkpoint's own DecisionEngine.predict — never the
semif mapping, whose different prompt would change the answers. Reuses the one
inference thread, bearer auth, MAX_QUEUE, VRAM cap and error envelope; no new
concurrency. 1..16 questions in ONE call (never chunked: Jev questions share a
prompt); images 422; over MAX_TOKENS 422 before the forward. /health advertises
the surface. Response 'model' is a string name@revision (JevBench's runner
hashes it; a dict broke its manifest step).

Acceptance on the live service (see acceptance/systemone-2026-09-30/): JevBench
v1.2.16 typesafe adapter over the 231 public items scores all 202/231, hard
83/111, with 0 changed answers across all 924 rows of the bench's own r1..r4;
controls 401/422x3 (token boundary proven at 7168 pass / 7169 refuse); GPU 1
per-process peak 9,866 MiB under the largest accepted request (budget 9,876);
/decide/shared positive control unchanged. 120 tests green.
2026-09-30 12:58:19 -07:00
vh 579b1f5d42 memory: Jev /v1/systemone tasked to infra-hermes; infra-ops audits 2026-09-30 12:39:16 -07:00
vh 2734cbf575 memory: losing Jev candidate weights deleted; intern-decision Jev-API status (model yes, service no) 2026-09-30 12:34:49 -07:00
vh 3c5f1ea803 docs(scriberr): slicer patch live on fv-ml1 as local-blackwell-a353078-slicer1; live GPU 1 peak 5,496 MiB
Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
2026-09-30 12:14:09 -07:00
vh ee3db68db1 feat(scriberr): overlap-and-stitch Parakeet slicer patch, rebuild script, bench
Carry patches/0001 on our Scriberr build (upstream a353078): adjacent
buffered chunks overlap by 4 s inside --chunk-len and hand over at a word
both chunks transcribed alike, instead of cutting at fixed marks with no
overlap. Pause-aware cutting is included as an opt-in (--pause-search);
it measured neutral once the stitch was right. The Go<->Python CLI and
JSON seam is unchanged.

Bench (4 recordings, 118 min, 3 cut placements each, against a no-cut
whole-file reference; metrics only, private audio stays on fv-ml1):
cuts with an error within +-3 s fall from 52% (93/179) to 22% (41/184)
against a 19% background; floor +-0.08. Positive control: upstream's
cutter +0.33 over background. A-vs-A byte-identical in-process and
across CLI processes. Peak GPU memory unchanged at 5,496 MiB (n=3).
Also found: Parakeet skips runs of >=10 words mid-chunk with any
slicer, upstream's included; not addressed here.

scripts/scriberr-rebuild clones a pinned upstream sha into a new
/opt/docker/src dir, git-apply-checks the patches, builds a distinct
tag, and checks embed, unit tests, the JSON seam (scriberr-seam-check.py)
and the memory budget on idle GPU 3. Deploy stays manual. The upstream
PR is prepared under patches/upstream-pr/ and not opened.
2026-09-30 12:10:32 -07:00
vh ab62644315 memory: scriberr pause-aware slicer build in flight 2026-09-30 10:53:03 -07:00
vh 128d1d847c semif: container removed after replacement by intern-decision; rollback is compose up -d 2026-09-30 09:49:30 -07:00
vh 1cf763a7b1 docs(intern-decision): live on fv-ml1 GPU 1; semif marked REPLACED
intern-decision deployed 0941 PT (cap 9.0 GiB, MAX_TOKENS 7168). Live acceptance: positive
control 240/259 and Wyrd 79/84, bit-identical to the bench (0/560 rows, Δp 0); negative control
10/122/14; largest accepted requests 200 with no 503; per-process 8,812 MiB at rest and 9,866 peak;
GPU 1 Free 15,442 before and 6,581 after (lowest 5,569 under load). Latency from nh3-dev:
21 criteria 114 ms, 16 over ~3,900 tokens 238 ms. semif README banner now REPLACED with the
rollback; fv-ml1 GPU 1 note updated.
2026-09-30 09:48:30 -07:00
vh 750675e391 feat(intern-decision): cap 9.0 GiB with MAX_TOKENS 7168, the largest call measured to fit
Both are required in compose because they are coupled: MAX_TOKENS is checked before the forward
pass, so an oversized call is a clear 422 instead of reaching the cap as a 503. Pre-deploy floor
is nvidia-smi Free >= 15,400 MiB on GPU 1 (card peak 9,876 + scriberr 5,496).
2026-09-30 09:40:08 -07:00
vh 66034cc69e scriberr: correct the GPU 1 budget — nvidia-smi Free is 15,442 MiB, not total−used; 70 MiB spare beside intern-decision at 9.0 GiB 2026-09-30 09:39:12 -07:00
vh a262477a61 feat(intern-decision): stack, DNS and GPU 3 acceptance for the SemIf replacement
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
2026-09-30 09:38:00 -07:00
vh 21d16d7ad8 fix(intern-decision-serve): code-review fixes
- a failure while building the response (a non-finite number included) is a 500 inside the
  envelope, never a 422 or a render crash outside it
- an engine ValueError keeps its message but is released and raised unchained, like an OOM
- the prompt is built (and the model's own validation run) before the forward
- the row cap is counted before any ordering is built
- /health reads a device name cached at load, so it makes no driver call off the inference thread
2026-09-30 09:30:40 -07:00
vh 618390c5fa docs(scriberr): Parakeet memory table, local-attention OOM, scripts are rewritten from the embedded copy 2026-09-30 09:21:51 -07:00
vh f21369e4ac fix(intern-decision-serve): load and score on one dedicated inference thread
torch keeps CUDA state per host thread (cuBLAS handles and workspaces), partly outside the
per-process VRAM cap. Scoring on anyio's threadpool let 40 threads each create it: measured on
fv-ml1 GPU 3, +252 MiB outside the cap and +326 MiB inside, which pushed the process past the
10,300 MiB GPU 1 budget. Load, warm-up and every call now run on the same single thread.
2026-09-30 09:18:46 -07:00
vh f7415db5c9 fix(intern-decision-serve): register the checkpoint's inference module before executing it
Its dataclasses use postponed annotations and look their module up in sys.modules while the
class is built; importing by file path without registering it failed startup (closed).
2026-09-30 09:08:19 -07:00
vh cc211ef10c scripts: wt-h2-count.py — worldtree-dev's U11b H2 gate check (verbatim); rehearsed on copies, 0 on both 2026-09-30 09:07:50 -07:00
vh 5bbf0aaeba feat(intern-decision-serve): Intern-Decision-4B behind semif-serve's HTTP surface
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
2026-09-30 09:04:39 -07:00
vh bb806e3596 memory: U11b step-5 auto-trigger (3 PASS) is mine; semif->intern-decision in flight; scriberr GPU budget 2026-09-30 09:03:21 -07:00
vh 0176ec0a5b scriberr: fit GPU 1 beside intern-decision — 120 s Parakeet slices + expandable_segments
Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s
9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with
expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the
recreated container's own env. Transcripts: 95.8% word-sequence similarity
vs 300 s, diffs mostly casing/punctuation.
2026-09-30 09:02:45 -07:00
vh d9bbaa07b2 memory: Jev bench done — Intern-Decision-4B is the SemIf replacement candidate 2026-09-30 05:01:30 -07:00
vh 475d6d6bcb docs: Jev candidate bench vs SemIf (fv-ml1 GPU 3) — Intern-Decision-4B is the replacement if SemIf is displaced
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
2026-09-30 05:00:45 -07:00
vh 9a6ac59da7 memory: U11b legacy-memory archive done (dedicated restic repo, drill passed); destroy-by 2026-10-30 runbook 2026-09-30 03:01:54 -07:00
vh 37d0b34682 memory: U11 daily off batches live (infra-hermes), first PASS; U11b /embed usage + legacy inventory + archive plan 2026-09-30 02:24:28 -07:00
vh aee8de3395 wt-memory-gate-batch: busy is exit 4 (retry later), and PASS is reported too
Both verdicts now build the n for the U11b data-deletion gate (three
consecutive PASS batches at off). First real run 20260930T090608Z via the
detached-worktree path: PASS, user median 0.83.
2026-09-30 02:24:05 -07:00
vh 6dd9ed2964 memory: U11a overnight log sweep clean but traffic-free; re-sweep TODO 2026-09-30 01:55:01 -07:00
vh 21e064b588 memory: scriberr 35-min retry succeeded with SemIf offline 2026-09-30 01:39:07 -07:00
vh 9daf43683d semif: offline by operator ruling — scriberr needs GPU 1 headroom
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
2026-09-30 01:36:18 -07:00
vh 5186b3dc9a scripts: wt-memory-gate-batch — one Worldtree U8 gate batch against demo's deployed sha
For infra-hermes's daily batch through the U11 off window (operator ruling
2026-09-29 2340; mode revised to off 2026-09-30). Enforces worldtree-dev's
terms: flock plus a refusal while any harness batch runs, code under test =
demo's deployed sha (a detached worktree when the main tree has moved; refuses
on pyproject/uv.lock/packages drift the shared venv cannot honour), no git
commit, no retry. Exit 0 PASS / 1 FAIL / 2 harness error / 3 refused.
2026-09-30 01:26:46 -07:00
vh 060fd92e4e memory: Worldtree U11a — personal flipped to legacy off and verified 2026-09-30 01:21:13 -07:00
vh 9905b76404 memory: Worldtree U11a — Prime ruled legacy off; demo flipped via config repo and verified; personal pending 2026-09-30 01:17:23 -07:00