Commit Graph
1507 Commits
Author SHA1 Message Date
vh e73adfe7ae dns+memory: nh3-pve-2-amt 10.100.250.63 (UDM port 5 on nh3-mgmt, reservation) 2026-10-02 14:43:16 -07:00
vh 59c9c0715a memory: MS-03 roles — esh-dev inherits nh3-dev's sessions; nh3-pve-2 purpose TBD by design 2026-10-02 09:17:39 -07:00
vh 536230170e memory: one MS-03 deploys as nh3-pve-2 (AMT on DHCP for phone-home; addressing plan) 2026-10-02 09:10:31 -07:00
vh 9dba6baccf homepage: NH3-PVE-AMT card opens MeshCentral (AMT now phones home; LAN management is off) 2026-10-02 09:06:49 -07:00
vh 1a2d763de4 ops(nh3-pve): AMT phones home to MeshCentral (CIRA); AMT on DHCP; LAN management now dark by design 2026-10-02 09:06:24 -07:00
vh 9bcb9417fb feat(amt): amt-cira-setup.py — configure Intel AMT phone-home (CIRA) to MeshCentral over WS-Man
MeshCentral only pushes CIRA through an on-host agent, so this mirrors its
amtmanager.js sequence directly over the LAN: trust MeshCentral's root,
add the MPS (username = 16-char meshid prefix), a periodic policy,
BIOS+OS user-initiated connections and a random environment-detection
domain. Idempotent, with a read-only state report. Enumerate uses
Enumerate+Pull (AMT 16 ignores OptimizeEnumeration; caught by a
positive control that first read 0 instances of a class that has one).

Applied to nh3-pve's AMT 2026-10-02 alongside MeshCentral mpsPass and an
ana-gw VIP/policy for 4433; the AMT does not dial out yet (static IP).
2026-10-02 08:50:47 -07:00
vh 61ba384d28 memory: demo fix-forward 7ab6ae40 verified; deploy-gate gap filed as Worldtree #423 2026-10-02 08:24:33 -07:00
vh 1219cfaca9 memory: demo rolled back from broken c2d87263 to fa8bc51c; b193 retired-key list for demo 2026-10-02 08:17:24 -07:00
vh 8b81579bca ops(pfi-tacticalrmm): MeshCentral to hybrid mode; nh3-pve AMT added and connected 2026-10-02 08:14:19 -07:00
vh a967bb95a6 docs(pfi-tacticalrmm): MeshCentral facts — WAN-only mode silently drops AMT adds; CLI access via the vaulted login token 2026-10-02 08:09:02 -07:00
vh cd30592e55 fix(wt-memory-gate-batch): pass --legacy-mode only when the deployed harness has it
b193 (U11b) retired the legacy plane and removed --legacy-mode from the
harness; passing it there fails the daily batch. The wrapper now greps
core/memory_acceptance at the deployed sha and omits the flag when it is
gone, so it is correct on both sides of the b192 -> b193 deploy (checked:
7a83f2f -> passes it, c2d87263 -> omits it).
2026-10-02 07:59:03 -07:00
vh 1cc8ad3292 memory: U11b step 5 done — live legacy memory deleted on demo and personal (H2 0), stamp sent 2026-10-02 07:54:50 -07:00
vh b0789f4b29 memory: U11b — gate-batch wrapper drops --legacy-mode at b193 2026-10-02 07:50:38 -07:00
vh 6e15f610bc memory: U11b gate streak 3 of 3; step 5 held for Prime's direct go 2026-10-02 07:50:12 -07:00
vh 59fb6b77a8 ops(nh3-dev): seat healthz watchdog — the 5-day pull-only outage must page, not age
althing-seat-daemon took an outside SIGTERM on 2026-09-27 and exited 0, so
Restart=on-failure never revived it; hermes-gateway sat pull-only for five
days with a yellow verdict nobody was watching. Fixed the death mode with a
Restart=always drop-in (also on the jekyll twin, same latent bug) and added
this watchdog for every other death mode: 5-min timer, alert on the second
consecutive non-200 (~10 min sustained), re-alert at most hourly, recovery
mail on green, exit 2 distinguishes a broken alarm wire from a down seat
(backup-freshness precedent). Lifecycle exercised end-to-end against a dead
port before install.
2026-10-02 07:47:17 -07:00
vh 902e16630f feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)
Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
2026-10-02 07:42:45 -07:00
vh 49c8fdf383 memory: nh3-dev snapshot waits for a natural reboot (Prime); post-boot checks listed 2026-10-02 07:36:15 -07:00
vh 6ec76bf35c ops(nh3-dev): grow root into the 378 GB disk online; swap moved to /swapfile
Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85%
twice in 14 h. The swap partition sat right after sda1 and blocked growth,
so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows
sda1 in place (start sector unchanged) and ext4 online, and sets initramfs
RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%.
2026-10-02 07:33:36 -07:00
vh 01c356e7fb docs(ana-docker): record the RustDesk server key and where the client setup steps live 2026-10-01 15:31:32 -07:00
vh abcf4c3166 fix(scriberr): serve over HTTPS via the fleet TLS caddy so the browser recorder works
The in-browser recorder calls getUserMedia, which browsers refuse on http://
origins, so it sat at "Initializing recorder...". scriberr.nh3.phasefinal.com
is now fronted by the fleet TLS caddy on nh3-dev (wildcard cert) and added to
Scriberr's ALLOWED_ORIGINS; the Homepage link points at it. The plain
http://10.251.50.54:8080 URL keeps working except for recording.
2026-10-01 13:07:18 -07:00
vh 474f78811c memory: U11b — writer/reader default to enabled at b193 (no instance action) 2026-10-01 12:35:57 -07:00
vh f37a50c11a memory: U11b C5 — memory_extractor and memory.escalation cleared to retire after b193 2026-10-01 11:25:01 -07:00
vh aa440708c4 memory: U11b C5 post-b193 config retirements listed; memory_extractor held open (escalation.model_role) 2026-10-01 11:23:32 -07:00
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00
vh e5536784e0 memory: U11b gate verdicts arrive from hermes-gateway; reply to infra-hermes 2026-10-01 07:49:49 -07:00
vh 23cedb3259 memory: Worldtree U11b gate streak 2 of 3 (20261001T143101Z PASS) 2026-10-01 07:49:28 -07:00
vh de11e00dfb fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
2026-10-01 05:59:07 -07:00
vh 9af5a16d07 memory: snapshot — dev-backup retention broken since 07-18 (log rotated, errors counted; Prime to rule); gen-small KV + graphs-off left as is 2026-10-01 05:29:27 -07:00
vh 0d9bb109b8 memory: gen-small OOM fixed by parakeet-nemo 0.1.1 (cache return + hard cap); Gitea webhook allow-list lesson 2026-10-01 04:42:13 -07:00
vh 963f9ed8c0 fix(parakeet-nemo): return the window cache and cap the process (nemo-0.1.1)
The gen-small EngineCore OOM (04:21 PT): our parked 3,582 MiB window cache left
no room for vLLM's runtime workspace. Seat-side fix, three controls:

- windowed path wraps every window in torch.cuda.empty_cache(), so the seat
  returns to ~2,108 MiB rest after a 12-min file instead of parking at the
  peak (measured: peak 3,028 MiB during, rest after, restarts=0);
- MEM_CAP_MIB=3840 hard set_per_process_memory_fraction: over-cap requests
  answer 503 with the seat alive (proved at cap=2000), so the failure lands
  on us, never on a neighbour;
- CUDA_GRAPHS=0: the graph decoder pins cache blocks that empty_cache must
  free (illegal-memory-access wedge when both were on first try). Cost:
  12-min file 3.0 s vs 1.2 s, short bins 35-62 ms vs 33-42 ms -- still 4-15x
  under the sherpa seat.

Measured, not computed: gen-small moved ZERO from 36,116 MiB across three
realistic requests (1,351 in / ~180 out) -- its workspace lands at engine
init; the growth window is restart-relative, matching infra-ops's observation.
WINDOW_S is now a real compose tunable. README memory section rewritten.
2026-10-01 04:39:39 -07:00
vh a816443acd memory: arbo webhook repointed off the retired wg0 IP 2026-10-01 04:30:03 -07:00
vh f83e35b94b memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived 2026-10-01 04:24:44 -07:00
vh 6b66207b0c scriberr: no upstream contribution (Prime) — drop prepared PR text; patches carried locally 2026-10-01 04:20:17 -07:00
vh 03826d2029 docs: gen-small .env.example records util 0.33 + boot-check rule; parakeet-nemo compose note corrected (slack, not KV; steady state 3,582 MiB) 2026-10-01 01:38:42 -07:00
vh cb28d8c951 memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state 2026-10-01 01:37:49 -07:00
vh 392660bd5f docs(parakeet-nemo): steady-state 3,582 MiB and the boot-check arithmetic rule (audit findings) 2026-10-01 01:37:47 -07:00
vh 41d2014df9 docs(parakeet): mark the sherpa seat SUPERSEDED; keep it as the rollback build source 2026-10-01 01:32:57 -07:00
vh de6ea32f34 feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
2026-10-01 01:32:48 -07:00
vh dbd583d6ca memory: irv-ml1 storetank reclaim closed (Prime via comfy-dev); symlink-target lesson 2026-10-01 00:11:05 -07:00
vh aa7eeff445 memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits 2026-09-30 23:54:54 -07:00
vh 1ae324d576 memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived 2026-09-30 23:52:27 -07:00
vh 212b736836 memory: parakeet seat A/B done — runtime is the bottleneck; unified-en NeMo bf16 wins speed+WER; seat defects 2026-09-30 18:53:04 -07:00
vh b38ec6dea0 docs(parakeet): seat A/B - clock times in Pacific military time 2026-09-30 18:51:59 -07:00
vh a6c1d3c454 docs(parakeet): seat A/B vs parakeet-unified-en-0.6b - latency is the int8-on-CPU runtime; unified wins WER
A/B of the live STT seat (fv-ml1 GPU 0, sherpa-onnx int8 v3) against
nvidia/parakeet-unified-en-0.6b, measured on GPU 3 with the seat's own image,
k2-fsa's published unified int8 export, fp32/fp16 exports made with k2-fsa's
recipe, v2 int8, and NeMo 3.0.0 (fp32, bf16 autocast, bf16 weights).

- Seat int8 graph runs on one CPU thread (cpu/wall 1.00, GPU 2-9%).
- unified-en under NeMo: -121/-234/-530 ms vs the seat at 1-3/3-8/8-20 s
  (paired, n=120/bin; floor <=6 ms; +50 ms positive control reads +52-54).
- unified-en WER lower in every runtime: -0.7 pp clean, -1.5 pp other,
  -3.2 to -4.4 pp AMI (paired CIs exclude 0).
- Seat defects found: hard 400 s input ceiling (HTTP 500), truncation after
  a quiet 1.5 s pause, and severe long-window dropouts (int8 v3 only).
- B-bf16w needs +0.8 to +1.5 GB over the seat's 1,690 MiB on GPU 0.

Raw requests, hypotheses, manifests and the full harness under
services/parakeet-ab-2026-09-30/. No deploy; live seat untouched apart
from 240 light test requests.
2026-09-30 18:51:44 -07:00
vh 11174ffea1 memory: GPU 0 room can come from trimming gen-small KV (Prime); donor analysis 2026-09-30 16:08:10 -07:00
vh aa17f64bd1 memory: warm-up cache audit passed; parakeet seat A/B in flight 2026-09-30 16:05:56 -07:00
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00
vh f65e27b08f feat(intern-decision): persist the triton autotune cache across recreates
The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable
layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image
owned by 10001 so the named volume intern-decision_triton-cache inherits a
writable mount point, and compose mounts it.

Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After
one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst
2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s.

scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from
/health, two-point live calibration of the tokenizer's linear token model (a
single probe overcorrects and the aim oscillates around the bucket edge),
per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE
CHANGE only; the volume carries ordinary recreates (measured: force-recreate,
then a warmed 32k call answered in 2.11 s).

Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU
1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README
caveat); /decide answers. Artifacts in the acceptance dir.
2026-09-30 16:02:16 -07:00
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00
vh 1189adbf18 memory: intern-decision warm-up cache tasked to infra-hermes 2026-09-30 15:44:45 -07:00