Commit Graph
583 Commits
Author SHA1 Message Date
vh aa440708c4 memory: U11b C5 post-b193 config retirements listed; memory_extractor held open (escalation.model_role) 2026-10-01 11:23:32 -07:00
vh f44280e6b2 feat(mia): one-shot Make-It-Animatable v2 auto-rigger on fv-ml1 GPU 3 (scripts/mia-run)
Image local/mia:0.1.0 built from stacks/mia: MIA v2 @ bbd8b158 (MIT) with
its pinned submodules, dread-dev's proven Python lock with the torch family
swapped to cu129, and a driver adapted from dread-dev's run_mia.py that
seeds every mesh (fix_random + trimesh's module RNG) and writes
weights_effective into the npz. Weights stay in fv-ml1's shared HF cache
at pinned revisions, mounted read-only.

scripts/mia-run mirrors blender-run: --job DIR is shipped to
fv-ml1:/tank/mia/jobs, one docker run --rm rigs every mesh, out/ comes back.

Acceptance on the four Dread Naught characters: 3.9-4.7 s a mesh (median
of 3) plus 12.7 s model load, peak 3,394 MiB; seeded runs bit-identical
across rotated mesh order; GPU-vs-CPU distances the same size as sampling
noise, with an unseeded GPU run as the positive control.
2026-10-01 09:45:22 -07:00
vh e5536784e0 memory: U11b gate verdicts arrive from hermes-gateway; reply to infra-hermes 2026-10-01 07:49:49 -07:00
vh 23cedb3259 memory: Worldtree U11b gate streak 2 of 3 (20261001T143101Z PASS) 2026-10-01 07:49:28 -07:00
vh de11e00dfb fix(dev-backup): chmod before prune so retention actually deletes; fail the unit on a bad prune
rsync -a copies a read-only source dir (0555) as read-only, so the hourly
prune's rm -rf could not unlink inside it. From 2026-07-18 every pruned
snapshot was left as a 22-entry husk while the run still logged OK. The
1,740 husks on nh3-nas were removed (0 errors; 48 full snapshots kept).

The prune now runs chmod -R u+w before rm -rf, logs its error count and
the number of snapshots kept, and exits 3 on a failed prune (exit 1 on a
failed rsync) so the systemd unit shows failed instead of passing.
2026-10-01 05:59:07 -07:00
vh 9af5a16d07 memory: snapshot — dev-backup retention broken since 07-18 (log rotated, errors counted; Prime to rule); gen-small KV + graphs-off left as is 2026-10-01 05:29:27 -07:00
vh 0d9bb109b8 memory: gen-small OOM fixed by parakeet-nemo 0.1.1 (cache return + hard cap); Gitea webhook allow-list lesson 2026-10-01 04:42:13 -07:00
vh a816443acd memory: arbo webhook repointed off the retired wg0 IP 2026-10-01 04:30:03 -07:00
vh f83e35b94b memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived 2026-10-01 04:24:44 -07:00
vh cb28d8c951 memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state 2026-10-01 01:37:49 -07:00
vh dbd583d6ca memory: irv-ml1 storetank reclaim closed (Prime via comfy-dev); symlink-target lesson 2026-10-01 00:11:05 -07:00
vh aa7eeff445 memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits 2026-09-30 23:54:54 -07:00
vh 1ae324d576 memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived 2026-09-30 23:52:27 -07:00
vh 212b736836 memory: parakeet seat A/B done — runtime is the bottleneck; unified-en NeMo bf16 wins speed+WER; seat defects 2026-09-30 18:53:04 -07:00
vh 11174ffea1 memory: GPU 0 room can come from trimming gen-small KV (Prime); donor analysis 2026-09-30 16:08:10 -07:00
vh aa17f64bd1 memory: warm-up cache audit passed; parakeet seat A/B in flight 2026-09-30 16:05:56 -07:00
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00
vh 1189adbf18 memory: intern-decision warm-up cache tasked to infra-hermes 2026-09-30 15:44:45 -07:00
vh af450f4ef7 docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up 2026-09-30 15:41:55 -07:00
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
vh e3dbb08d84 memory: true Jev is text-only (official docs); our real Jev gap is context (7,168 vs 64k tokens) 2026-09-30 13:09:44 -07:00
vh 1faadb45f6 memory: intern-decision 0.1.2 live, re-audit passed 2026-09-30 13:06:53 -07:00
vh da50696fe8 memory: intern-decision /v1/systemone live (0.1.1), infra-ops audit passed; 2 low findings to infra-hermes 2026-09-30 13:01:05 -07:00
vh 579b1f5d42 memory: Jev /v1/systemone tasked to infra-hermes; infra-ops audits 2026-09-30 12:39:16 -07:00
vh 2734cbf575 memory: losing Jev candidate weights deleted; intern-decision Jev-API status (model yes, service no) 2026-09-30 12:34:49 -07:00
vh 3c5f1ea803 docs(scriberr): slicer patch live on fv-ml1 as local-blackwell-a353078-slicer1; live GPU 1 peak 5,496 MiB
Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
2026-09-30 12:14:09 -07:00
vh ab62644315 memory: scriberr pause-aware slicer build in flight 2026-09-30 10:53:03 -07:00
vh 128d1d847c semif: container removed after replacement by intern-decision; rollback is compose up -d 2026-09-30 09:49:30 -07:00
vh 66034cc69e scriberr: correct the GPU 1 budget — nvidia-smi Free is 15,442 MiB, not total−used; 70 MiB spare beside intern-decision at 9.0 GiB 2026-09-30 09:39:12 -07:00
vh cc211ef10c scripts: wt-h2-count.py — worldtree-dev's U11b H2 gate check (verbatim); rehearsed on copies, 0 on both 2026-09-30 09:07:50 -07:00
vh bb806e3596 memory: U11b step-5 auto-trigger (3 PASS) is mine; semif->intern-decision in flight; scriberr GPU budget 2026-09-30 09:03:21 -07:00
vh d9bbaa07b2 memory: Jev bench done — Intern-Decision-4B is the SemIf replacement candidate 2026-09-30 05:01:30 -07:00
vh 9a6ac59da7 memory: U11b legacy-memory archive done (dedicated restic repo, drill passed); destroy-by 2026-10-30 runbook 2026-09-30 03:01:54 -07:00
vh 37d0b34682 memory: U11 daily off batches live (infra-hermes), first PASS; U11b /embed usage + legacy inventory + archive plan 2026-09-30 02:24:28 -07:00
vh 6dd9ed2964 memory: U11a overnight log sweep clean but traffic-free; re-sweep TODO 2026-09-30 01:55:01 -07:00
vh 21e064b588 memory: scriberr 35-min retry succeeded with SemIf offline 2026-09-30 01:39:07 -07:00
vh 9daf43683d semif: offline by operator ruling — scriberr needs GPU 1 headroom
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
2026-09-30 01:36:18 -07:00
vh 060fd92e4e memory: Worldtree U11a — personal flipped to legacy off and verified 2026-09-30 01:21:13 -07:00
vh 9905b76404 memory: Worldtree U11a — Prime ruled legacy off; demo flipped via config repo and verified; personal pending 2026-09-30 01:17:23 -07:00
vh f9fae53599 memory: snapshot — Worldtree U10 done, U11a prepped, U8 wrapper for infra-hermes in progress; Blender extensions live; Bonsai spike closed; 3 entries archived 2026-09-29 23:42:03 -07:00
vh 274b8175c0 memory: Worldtree U11a prepped (staged config on demo, U8 window harness facts) 2026-09-29 23:33:24 -07:00
vh 47d8fd6139 memory: Worldtree U10 backfill committed on personal (797 filed) 2026-09-29 10:07:45 -07:00
vh 40427b4822 memory: U10 personal presync; demo U9 forget-policy gap fixed; pinned chroma in container layer 2026-09-29 08:55:16 -07:00
vh b7b6e01d6e memory: Worldtree U10 backfill committed on demo; memory_tagger config sync needed on personal 2026-09-28 23:52:05 -07:00
vh ea5d8dfd4e docs(blender): desktop Blender + MCP registration per working session, not per task (Prime 2026-09-28) 2026-09-28 16:36:44 -07:00
vh 2296fbba12 memory: Bonsai fork pin has a second copy on the smithy NAS 2026-09-28 15:44:44 -07:00
vh 60e6ed5b4e docs(fv-ml1): /tank/aimodels/mlx and the PROVENANCE-beside convention; Bonsai weights acquired 2026-09-28 15:43:52 -07:00
vh fd20183cbb feat(blender): extension set live in the GUI; MCP acceptance 9/9, probe made safe-mode compliant 2026-09-28 15:08:22 -07:00
vh 056433555f memory: Bonsai MMVQ-threshold follow-up (N=8 1.06x -> 1.21x); build image removed 2026-09-28 14:33:39 -07:00