Commit Graph
605 Commits
Author SHA1 Message Date
vh 3a91eff2e6 memory: ESH tank scrub repairing the Aug 20 replacement's unrepaired blocks 2026-10-02 23:33:07 -07:00
vh 55004090d1 ops(esh-pve-2): register host, AMT phoning home to MeshCentral; note MeshCentral first-CIRA crash race 2026-10-02 22:25:37 -07:00
vh 6315ce38d5 memory: ESH tank verified residual, cleared, full scrub running 2026-10-02 19:26:30 -07:00
vh 62254023f3 docs: esh-pve-cluster PVE 8->9 upgrade plan (not executed); ESH tank DEGRADED finding 2026-10-02 19:20:28 -07:00
vh 5439d5fdc1 albok-service: 0.1.1 (health fix) on nh3-docker 2026-10-02 18:04:57 -07:00
vh d3958655c1 feat(albok-service): deploy the fleet knowledgebase service on nh3-docker
albok-service 0.1.0 (vh/albok 0b37431), image pfi/albok-service pinned by
digest, published on host port 8392 because 8390 is the post office.
Host prep playbook creates the fixed ids (albok 1500, albok-read 1510,
albok-personal 1511) and the local store/private roots; the container
gets a mounted /etc/group and group_add so the service can resolve and
chgrp its wing dirs. The config carries a LiteLLM key scoped to
qwen3-embedding and lives outside the deploy-synced conf dir. DNS name
albok.nh3.internal.
2026-10-02 17:56:34 -07:00
vh 9c233ca255 memory: nh3-pve-2 end-of-day state and on-site reinstall checklist 2026-10-02 17:40:50 -07:00
vh 34659ae9e7 ops(nh3-pve-2): AMT 21 configured (KVM, no-consent, listener) and phoning home to MeshCentral 2026-10-02 14:54:56 -07:00
vh e73adfe7ae dns+memory: nh3-pve-2-amt 10.100.250.63 (UDM port 5 on nh3-mgmt, reservation) 2026-10-02 14:43:16 -07:00
vh 59c9c0715a memory: MS-03 roles — esh-dev inherits nh3-dev's sessions; nh3-pve-2 purpose TBD by design 2026-10-02 09:17:39 -07:00
vh 536230170e memory: one MS-03 deploys as nh3-pve-2 (AMT on DHCP for phone-home; addressing plan) 2026-10-02 09:10:31 -07:00
vh 1a2d763de4 ops(nh3-pve): AMT phones home to MeshCentral (CIRA); AMT on DHCP; LAN management now dark by design 2026-10-02 09:06:24 -07:00
vh 61ba384d28 memory: demo fix-forward 7ab6ae40 verified; deploy-gate gap filed as Worldtree #423 2026-10-02 08:24:33 -07:00
vh 1219cfaca9 memory: demo rolled back from broken c2d87263 to fa8bc51c; b193 retired-key list for demo 2026-10-02 08:17:24 -07:00
vh 8b81579bca ops(pfi-tacticalrmm): MeshCentral to hybrid mode; nh3-pve AMT added and connected 2026-10-02 08:14:19 -07:00
vh cd30592e55 fix(wt-memory-gate-batch): pass --legacy-mode only when the deployed harness has it
b193 (U11b) retired the legacy plane and removed --legacy-mode from the
harness; passing it there fails the daily batch. The wrapper now greps
core/memory_acceptance at the deployed sha and omits the flag when it is
gone, so it is correct on both sides of the b192 -> b193 deploy (checked:
7a83f2f -> passes it, c2d87263 -> omits it).
2026-10-02 07:59:03 -07:00
vh 1cc8ad3292 memory: U11b step 5 done — live legacy memory deleted on demo and personal (H2 0), stamp sent 2026-10-02 07:54:50 -07:00
vh 6e15f610bc memory: U11b gate streak 3 of 3; step 5 held for Prime's direct go 2026-10-02 07:50:12 -07:00
vh 902e16630f feat(dev-backup): add daily and weekly retention (48 hourly + 30 daily + 12 weekly)
Prime's ruling 2026-10-02. retention.py picks the snapshots to delete:
the newest 48, plus the newest of each of the last 30 days and of each of
the last 12 ISO weeks, counting only days and weeks that have snapshots.
Names that are not exactly YYYY-MM-DD_HHMM are never selected, and the
NAS side refuses any path outside that pattern. The unit fails unless
the number kept equals the number expected. Live run: deleted 1, 0
errors, 48 kept as expected.
2026-10-02 07:42:45 -07:00
vh 49c8fdf383 memory: nh3-dev snapshot waits for a natural reboot (Prime); post-boot checks listed 2026-10-02 07:36:15 -07:00
vh 6ec76bf35c ops(nh3-dev): grow root into the 378 GB disk online; swap moved to /swapfile
Prime grew VM 102's scsi0 from 250 to 378 GB after the root alert hit 85%
twice in 14 h. The swap partition sat right after sda1 and blocked growth,
so the playbook moves swap to a 4 GB /swapfile, deletes sda5/sda2, grows
sda1 in place (start sector unchanged) and ext4 online, and sets initramfs
RESUME=none so boots don't wait for the vanished swap. Root is 372 GB, 58%.
2026-10-02 07:33:36 -07:00
vh abcf4c3166 fix(scriberr): serve over HTTPS via the fleet TLS caddy so the browser recorder works
The in-browser recorder calls getUserMedia, which browsers refuse on http://
origins, so it sat at "Initializing recorder...". scriberr.nh3.phasefinal.com
is now fronted by the fleet TLS caddy on nh3-dev (wildcard cert) and added to
Scriberr's ALLOWED_ORIGINS; the Homepage link points at it. The plain
http://10.251.50.54:8080 URL keeps working except for recording.
2026-10-01 13:07:18 -07:00
vh aa440708c4 memory: U11b C5 post-b193 config retirements listed; memory_extractor held open (escalation.model_role) 2026-10-01 11:23:32 -07:00
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00
vh e5536784e0 memory: U11b gate verdicts arrive from hermes-gateway; reply to infra-hermes 2026-10-01 07:49:49 -07:00
vh 23cedb3259 memory: Worldtree U11b gate streak 2 of 3 (20261001T143101Z PASS) 2026-10-01 07:49:28 -07:00
vh de11e00dfb fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
2026-10-01 05:59:07 -07:00
vh 9af5a16d07 memory: snapshot — dev-backup retention broken since 07-18 (log rotated, errors counted; Prime to rule); gen-small KV + graphs-off left as is 2026-10-01 05:29:27 -07:00
vh 0d9bb109b8 memory: gen-small OOM fixed by parakeet-nemo 0.1.1 (cache return + hard cap); Gitea webhook allow-list lesson 2026-10-01 04:42:13 -07:00
vh a816443acd memory: arbo webhook repointed off the retired wg0 IP 2026-10-01 04:30:03 -07:00
vh f83e35b94b memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived 2026-10-01 04:24:44 -07:00
vh cb28d8c951 memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state 2026-10-01 01:37:49 -07:00
vh dbd583d6ca memory: irv-ml1 storetank reclaim closed (Prime via comfy-dev); symlink-target lesson 2026-10-01 00:11:05 -07:00
vh aa7eeff445 memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits 2026-09-30 23:54:54 -07:00
vh 1ae324d576 memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived 2026-09-30 23:52:27 -07:00
vh 212b736836 memory: parakeet seat A/B done — runtime is the bottleneck; unified-en NeMo bf16 wins speed+WER; seat defects 2026-09-30 18:53:04 -07:00
vh 11174ffea1 memory: GPU 0 room can come from trimming gen-small KV (Prime); donor analysis 2026-09-30 16:08:10 -07:00
vh aa17f64bd1 memory: warm-up cache audit passed; parakeet seat A/B in flight 2026-09-30 16:05:56 -07:00
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00
vh 1189adbf18 memory: intern-decision warm-up cache tasked to infra-hermes 2026-09-30 15:44:45 -07:00
vh af450f4ef7 docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up 2026-09-30 15:41:55 -07:00
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
vh e3dbb08d84 memory: true Jev is text-only (official docs); our real Jev gap is context (7,168 vs 64k tokens) 2026-09-30 13:09:44 -07:00
vh 1faadb45f6 memory: intern-decision 0.1.2 live, re-audit passed 2026-09-30 13:06:53 -07:00
vh da50696fe8 memory: intern-decision /v1/systemone live (0.1.1), infra-ops audit passed; 2 low findings to infra-hermes 2026-09-30 13:01:05 -07:00
vh 579b1f5d42 memory: Jev /v1/systemone tasked to infra-hermes; infra-ops audits 2026-09-30 12:39:16 -07:00
vh 2734cbf575 memory: losing Jev candidate weights deleted; intern-decision Jev-API status (model yes, service no) 2026-09-30 12:34:49 -07:00
vh 3c5f1ea803 docs(scriberr): slicer patch live on fv-ml1 as local-blackwell-a353078-slicer1; live GPU 1 peak 5,496 MiB
Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
2026-09-30 12:14:09 -07:00
vh ab62644315 memory: scriberr pause-aware slicer build in flight 2026-09-30 10:53:03 -07:00