Commit Graph
577 Commits
Author SHA1 Message Date
vh 0d9bb109b8 memory: gen-small OOM fixed by parakeet-nemo 0.1.1 (cache return + hard cap); Gitea webhook allow-list lesson 2026-10-01 04:42:13 -07:00
vh a816443acd memory: arbo webhook repointed off the retired wg0 IP 2026-10-01 04:30:03 -07:00
vh f83e35b94b memory: snapshot — speech seat live + gen-small OOM incident (mitigated, fix tasked); leftovers deleted; no Scriberr upstream; repos pushed; 32 entries archived 2026-10-01 04:24:44 -07:00
vh cb28d8c951 memory: parakeet speech seat switched to unified-en (NeMo) — live, audited; gen-small util 0.36, GPU 0 steady state 2026-10-01 01:37:49 -07:00
vh dbd583d6ca memory: irv-ml1 storetank reclaim closed (Prime via comfy-dev); symlink-target lesson 2026-10-01 00:11:05 -07:00
vh aa7eeff445 memory: parakeet seat switch tasked to infra-hermes (Prime); infra-ops audits 2026-09-30 23:54:54 -07:00
vh 1ae324d576 memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived 2026-09-30 23:52:27 -07:00
vh 212b736836 memory: parakeet seat A/B done — runtime is the bottleneck; unified-en NeMo bf16 wins speed+WER; seat defects 2026-09-30 18:53:04 -07:00
vh 11174ffea1 memory: GPU 0 room can come from trimming gen-small KV (Prime); donor analysis 2026-09-30 16:08:10 -07:00
vh aa17f64bd1 memory: warm-up cache audit passed; parakeet seat A/B in flight 2026-09-30 16:05:56 -07:00
vh 8c68bacf2e scriberr: carry patch 0002 (gap retry + PARAKEET_MODEL_PATH), live as dropout2
Prime: Scriberr gets the basic fix, v3 stays (no NeMo 3.0.0 surgery). 0002 moves
from proposed/ into the carried set; scriberr-rebuild now applies 0001+0002 by
default (suffix dropout2) and its memory budget becomes a 5,600 MiB regression
guard (Scriberr is on GPU 3). Live on fv-ml1 1602: scripts rewritten from the
patched embed, a 20-min file at 5,502 MiB with retried_gaps reported.
2026-09-30 16:03:07 -07:00
vh 38015a1977 docs(scriberr): Parakeet dropout investigation; proposed 0002 (gap retry + model path)
Prime's ask (via the coordinator): investigate the "Parakeet skips
stretches of speech" finding, including other Parakeet weights.
Investigation only; nothing deployed.

Against ground truth (official SCOTUS transcript, Gutenberg #38916) the
drops are real: production v3 loses 140 / 66 clean words per transcript on
the two public files and ~50 on each private one (Whisper-referenced,
Canary-confirmed; adjudicator 129/129 correct on the calibration). Cause:
the v2/v3 0.6B weights collapse deep inside long full-attention windows;
the encoder output is degraded, the audio alone transcribes fine, and
1.1B TDT/RNNT/CTC and CTC-0.6B never do it. Decoding (CUDA graphs, greedy
variants, max_symbols, beam), slice length, local attention, loudness,
resampling and a noise floor do not fix it. Controls: A-vs-A, silence
positive control (>=15 words 36/36), null control, bootstrap floor.

Proposed patch 0002 re-transcribes >=3 s stretches where the audio holds
speech but no word came out (-80 to -90 % lost words on all four
recordings, lower WER, no invented text, +10 MiB) and adds an explicit
PARAKEET_MODEL_PATH with the loaded model recorded in JSON and ModelUsed.
Reviewed at high effort, all findings fixed; built and tested as
scriberr:local-blackwell-a353078-dropout2, not deployed.

scriberr-rebuild: --patches takes DIR[:DIR...]; embeds and seam-checks
both Parakeet scripts (seam-check --standard for the short-audio one).
2026-09-30 15:48:02 -07:00
vh 1189adbf18 memory: intern-decision warm-up cache tasked to infra-hermes 2026-09-30 15:44:45 -07:00
vh af450f4ef7 docs(intern-decision): Triton warm state survives restart, lost on recreate; ~16 x 2048-token buckets, ~2 min full warm-up 2026-09-30 15:41:55 -07:00
vh 6b201e1d4a scriberr to fv-ml1 GPU 3 (on-demand, steps aside to irv-ml1 A6000); intern-decision 32k-token calls (cap 14.4 GiB)
Prime 2026-09-30: move scriberr to GPU 3 and extend the Jev endpoint to 32k tokens.
Scriberr holds 0 VRAM idle; verified a 20-min job on GPU 3 at 5,496 MiB. With GPU 1
freed, intern-decision's measured card peak at MAX_TOKENS=32768 is 15,220 MiB against
a 15,437 MiB budget (n=3, 1 and 16 questions); 32,769 tokens is refused 422 up front.
JevBench v1.2.16 via /v1/systemone unchanged: 202/231, 0 diffs vs the bench.
2026-09-30 13:35:32 -07:00
vh e3dbb08d84 memory: true Jev is text-only (official docs); our real Jev gap is context (7,168 vs 64k tokens) 2026-09-30 13:09:44 -07:00
vh 1faadb45f6 memory: intern-decision 0.1.2 live, re-audit passed 2026-09-30 13:06:53 -07:00
vh da50696fe8 memory: intern-decision /v1/systemone live (0.1.1), infra-ops audit passed; 2 low findings to infra-hermes 2026-09-30 13:01:05 -07:00
vh 579b1f5d42 memory: Jev /v1/systemone tasked to infra-hermes; infra-ops audits 2026-09-30 12:39:16 -07:00
vh 2734cbf575 memory: losing Jev candidate weights deleted; intern-decision Jev-API status (model yes, service no) 2026-09-30 12:34:49 -07:00
vh 3c5f1ea803 docs(scriberr): slicer patch live on fv-ml1 as local-blackwell-a353078-slicer1; live GPU 1 peak 5,496 MiB
Deployed 2026-09-30 1211 PT by pointing SCRIBERR_IMAGE at the patched tag (.env backed up as .env.bak-20260930-pre-slicer1; rollback is the unpatched scriberr:local-blackwell). PrepareEnvironment rewrote the env's parakeet_transcribe_buffered.py from the embed (sha256 matches the patched source). One live run on GPU 1 beside intern-decision peaked at 5,496 MiB. Memory records the open Parakeet mid-chunk dropout finding and the held upstream PR.
2026-09-30 12:14:09 -07:00
vh ab62644315 memory: scriberr pause-aware slicer build in flight 2026-09-30 10:53:03 -07:00
vh 128d1d847c semif: container removed after replacement by intern-decision; rollback is compose up -d 2026-09-30 09:49:30 -07:00
vh 66034cc69e scriberr: correct the GPU 1 budget — nvidia-smi Free is 15,442 MiB, not total−used; 70 MiB spare beside intern-decision at 9.0 GiB 2026-09-30 09:39:12 -07:00
vh cc211ef10c scripts: wt-h2-count.py — worldtree-dev's U11b H2 gate check (verbatim); rehearsed on copies, 0 on both 2026-09-30 09:07:50 -07:00
vh bb806e3596 memory: U11b step-5 auto-trigger (3 PASS) is mine; semif->intern-decision in flight; scriberr GPU budget 2026-09-30 09:03:21 -07:00
vh d9bbaa07b2 memory: Jev bench done — Intern-Decision-4B is the SemIf replacement candidate 2026-09-30 05:01:30 -07:00
vh 9a6ac59da7 memory: U11b legacy-memory archive done (dedicated restic repo, drill passed); destroy-by 2026-10-30 runbook 2026-09-30 03:01:54 -07:00
vh 37d0b34682 memory: U11 daily off batches live (infra-hermes), first PASS; U11b /embed usage + legacy inventory + archive plan 2026-09-30 02:24:28 -07:00
vh 6dd9ed2964 memory: U11a overnight log sweep clean but traffic-free; re-sweep TODO 2026-09-30 01:55:01 -07:00
vh 21e064b588 memory: scriberr 35-min retry succeeded with SemIf offline 2026-09-30 01:39:07 -07:00
vh 9daf43683d semif: offline by operator ruling — scriberr needs GPU 1 headroom
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
2026-09-30 01:36:18 -07:00
vh 060fd92e4e memory: Worldtree U11a — personal flipped to legacy off and verified 2026-09-30 01:21:13 -07:00
vh 9905b76404 memory: Worldtree U11a — Prime ruled legacy off; demo flipped via config repo and verified; personal pending 2026-09-30 01:17:23 -07:00
vh f9fae53599 memory: snapshot — Worldtree U10 done, U11a prepped, U8 wrapper for infra-hermes in progress; Blender extensions live; Bonsai spike closed; 3 entries archived 2026-09-29 23:42:03 -07:00
vh 274b8175c0 memory: Worldtree U11a prepped (staged config on demo, U8 window harness facts) 2026-09-29 23:33:24 -07:00
vh 47d8fd6139 memory: Worldtree U10 backfill committed on personal (797 filed) 2026-09-29 10:07:45 -07:00
vh 40427b4822 memory: U10 personal presync; demo U9 forget-policy gap fixed; pinned chroma in container layer 2026-09-29 08:55:16 -07:00
vh b7b6e01d6e memory: Worldtree U10 backfill committed on demo; memory_tagger config sync needed on personal 2026-09-28 23:52:05 -07:00
vh ea5d8dfd4e docs(blender): desktop Blender + MCP registration per working session, not per task (Prime 2026-09-28) 2026-09-28 16:36:44 -07:00
vh 2296fbba12 memory: Bonsai fork pin has a second copy on the smithy NAS 2026-09-28 15:44:44 -07:00
vh 60e6ed5b4e docs(fv-ml1): /tank/aimodels/mlx and the PROVENANCE-beside convention; Bonsai weights acquired 2026-09-28 15:43:52 -07:00
vh fd20183cbb feat(blender): extension set live in the GUI; MCP acceptance 9/9, probe made safe-mode compliant 2026-09-28 15:08:22 -07:00
vh 056433555f memory: Bonsai MMVQ-threshold follow-up (N=8 1.06x -> 1.21x); build image removed 2026-09-28 14:33:39 -07:00
vh 841f05de7a memory: Bonsai spike result on fv-ml1 GPU 3; blender-run --cpu 2026-09-28 14:23:21 -07:00
vh 59d436cc55 memory: first nightly restic run on repository-file verified all-green on all 8 hosts 2026-09-28 12:54:29 -07:00
vh 83dc497b40 feat(blender): pinned extension set in a read-only System repo, for the GUI and blender-run --extensions
Blender is now a mandatory stage in draupnir's pipeline (Prime, 2026-09-28), and draupnir asked
for eight add-ons from extensions.blender.org: SurfacePsycho 0.10.4, CAD Sketcher 0.32.1,
3D-Print Toolbox 1.4.1, STEP Importer 1.2.1, Bool Tool 2.1.0, LoopTools 4.7.7, MeasureIt 1.8.4,
3MF Import/Export 2.7.7.

- stacks/blender/extensions.lock pins each by version and archive sha256.
- scripts/blender-extensions sync builds fv-ml1:/tank/blender-extensions/5.2/system with Blender's
  own install-file, pre-warms and byte-compiles it, checks a read-only enable, then swaps it in.
  It refuses while the GUI or a blender-run job holds the old directory.
- conf/scripts/startup/fleet_extensions.py enables every package in the System repo: in a timer
  in the GUI (after the prefs load), and as --python ahead of the caller's args in
  blender-run --extensions (a failed enable exits 1 before the caller's script).
- It also patches SurfacePsycho's sp_overwrite_segment_selection from eval() to literal_eval():
  the eval walked past MCP safe mode (control: unpatched ran code, patched refuses).
- blender-run: --extensions (bind mounts via --mount so a missing source fails instead of being
  created); USER/LOGNAME set, which CAD Sketcher's getpass needs.
- compose.yaml mounts the repo read-only and the hook into the GUI container. NOT yet deployed.
- scripts/blender-probes/extensions_acceptance.py: one operator run per add-on, safe-mode
  compliant. Headless 8/9 online and with --network none; CAD Sketcher sketching is GUI-only.
  A Python audit hook saw no network/process events (positive control fired).
2026-09-28 12:50:57 -07:00
vh 1756b891c6 memory: snapshot — Zigbee2MQTT + HA-MQTT route fix, Blender on GPU 3 (MCP per task + blender-run), semif 0.1.4, Worldtree reward config live and pushed; 3 entries archived; next: 2026-09-28 restic freshness check 2026-09-28 09:57:41 -07:00
vh 6f0c480693 memory: Blender access stays on the shared fleet login; batch path blender-run (Prime) 2026-09-28 08:50:09 -07:00
vh 2b38cfd1b0 memory: worldtree-instance-configs pushed (Prime) 2026-09-27 15:05:15 -07:00