memory: snapshot — U11a off + U11b gate; SemIf→intern-decision (Jev, 32k); Scriberr GPU 3 + slicer + gap retry; Parakeet seat switch approved for next session; 26 entries archived
This commit is contained in:
+75
-137
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-30 ~0122 PT (Worldtree U11a: Prime ruled legacy OFF (not read_only); demo flipped 0115 and personal 0120 via the config repo, both verified. Prior: Blender extensions, Bonsai spike, phasefinal.com cleanup, U10 backfill.)_
|
||||
_Last updated: 2026-09-30 ~1800 PT (U11a legacy off on both Worldtree instances, U11b gate armed; SemIf → intern-decision with Jev /v1/systemone at 32k; Scriberr → GPU 3 with overlap slicer + gap retry; Parakeet seat switch to unified-en APPROVED, next session; 26 old entries archived.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -115,30 +115,69 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-09-30 ~0120 PT._
|
||||
_As of 2026-09-30 ~1800 PT._
|
||||
|
||||
### Worldtree memory-split: U10 done; U11a legacy OFF live on demo + personal (2026-09-30)
|
||||
### NEXT: switch the Parakeet speech seat (Prime 2026-09-30 ~1755: "reasonable terms, ship the switch")
|
||||
|
||||
- U10 legacy backfill DONE: demo 5 filed, personal 797 filed (mimir's 377 `dropped:too_large` are ONE interests record at its ceiling, re-fileable after consolidation; that is the U11 checkpoint's call). Pinned: Prime ruled "leave pinned".
|
||||
- **U11a: Prime ruled 2026-09-30 0100 (in worldtree-dev's session, thread `01M3RNRC8RE87AJ3M6XBNVHAYP`): legacy plane OFF, not read_only, on demo AND personal, as soon as the b192 image (64f79b38) lands.** The read_only window is skipped. Nothing is deleted; retirement of the old store is the next unit, and worldtree-dev announces it before any data goes. Accepted risk: skaldsong/wizard-v2 on personal, the only Tier-3 client, has unknown record-profile adoption, so its sessions stop being remembered until it adopts the profile. The create log says so per session with a WARNING `will not be remembered until the client adopts the record profile`.
|
||||
- **Config goes through the config repo, not a hand copy:** `~/development/worldtree-instance-configs` commit 63cf268 adds `memory.writer.enabled: true`, `memory.reader.enabled: true` and `memory.legacy.mode: "off"` to both instances' defaults.yaml. Commit 0a1387e captured the U10 memory_tagger and U9 forget-policy host edits that had never reached the repo. **Both commits are local; pushing them is Prime's call.**
|
||||
- ⚠ **`off` MUST be quoted.** `core/defaults.py` uses `yaml.safe_load`, which is YAML 1.1: a bare `off` parses as boolean False and the strict `LegacyMode` enum refuses the boot. Verified at 64f79b38 with Worldtree's own `load_cutover_config`. worldtree-dev's recipe had it bare.
|
||||
- **DEMO flipped at 0115 PT** with `deploy-wt-config deploy demo`. The gauge reads `worldtree_memory_legacy_mode{kind="off"} 1.0`, /health is 200, and the boot log is clean. The staged read_only file is deleted. Rollback: set legacy.mode to `"live"` in the repo and deploy it.
|
||||
- **PERSONAL flipped at 0120 PT** on 64f79b38a7fb, the same way. ⚠ Personal has **no /metrics**, because the #308 metrics block is demo-only, so the proof was the no-metrics one: an in-container `get_default→load_cutover_config→validate_cutover` gave off/writer/reader True, /health was 200, and the boot log matched the pre-flip one. Whether to add metrics to personal is still worldtree-dev's open call. Both instances are in sync with the repo.
|
||||
- Watch after each flip: any `legacy executor <kind> invoked with memory.legacy.mode=off; ignored` line goes to worldtree-dev with its stamp. A sweep at 0154 of both logs since their last start found **0 watched lines, but there was ~no traffic** (demo 0 real requests, personal 1), so that proves little. **TODO: re-sweep after the first real traffic** with `docker logs --since <StartedAt>`, grepping `legacy executor|Traceback|ERROR|will not be remembered`. A CI recreate drops the old container's log.
|
||||
- **Daily gate batches at off: LIVE, infra-hermes from 2026-10-01** (Prime's 2340 ruling; the plan was revised to `--legacy-mode off --label window-off` by worldtree-dev, thread `01M3RRW7WZ04GNR871HC1VKY4K`). The wrapper is `scripts/wt-memory-gate-batch`. It tests demo's deployed sha, from the detached worktree `~/development/Worldtree-gate` when worldtree-dev's tree has moved. Exit codes: 0 PASS, 1 FAIL, 2 error, 3 refused, 4 busy. **PASS and FAIL both go to worldtree-dev the same day**, because they build the n. **The U11b DATA deletion is gated on 3 consecutive PASS at off, or a Prime waiver**; code retirement is not gated. The controls were 080927Z FAIL and 082829Z FAIL (one flip each; the filing floor held; worldtree-dev calls it undecided). **Mine, 20260930T090608Z, PASSED (user median 0.83)**, so the count is 1 of 3 if the FAILs reset it. ⚠ It started by accident, from a test meant to be `--dry-run`, and matched the plan anyway.
|
||||
- **⚠ U11b STEP 5 IS MINE, AUTO-TRIGGERED (Prime ruled 0320 via worldtree-dev, thread `01M3SGEQDRQD7DWBVT4K73FAHP`):** once the THIRD CONSECUTIVE PASS at off lands (infra-hermes copies me on every verdict; 20260930T090608Z was 1 of 3; a FAIL restarts the count), I run the H2 check VERBATIM, per instance, right before its deletion: `docker exec -i <api> python - < scripts/wt-h2-count.py`. It is worldtree-dev's script, and exit 2 means STOP and send them the output. A rehearsal on `cp -a` copies on 09-30 read 0 on both (demo 4/12/5 rows, personal 0/6/2200). Then I delete on BOTH instances, using literal paths only: `agents/{forseti,lofn,mimir}/memory/<agent>.chroma` + `memory/context_promotion`. **I delete LIVE with the api running** (worldtree-dev, read at source on b192: at off, nothing holds those files open). ⚠ Any b192 restart re-creates an empty, schema-only `context_promotion/ledger.db` (+wal/shm). That is residue, not memory data; note it in the stamp and rm it after b193 is deployed. Then send worldtree-dev the stamp, because b193 ships after it. The 377 parked rows go with it. The archive stays under its rev 1.2 rules. After b193: remove the retired config keys at my pace.
|
||||
- **U11b prep (worldtree-dev asks, read-only, answered 0215):** /embed usage from the Skuld ledger (`phase_name='embed'`, which carries no caller identity): demo 436 calls, 07-29..08-31, none since; personal 5 calls, 07-30..08-05. ⚠ Demo's whole Skuld ledger has been idle since 09-14. Legacy data: demo 3 chroma ≈2.1M plus context_promotion 224K; personal ≈26M plus 6.5M (1,413 JSONL + ledger.db). Both are in the `*_worldtree-state` volumes. **ARCHIVE DONE 2026-09-30 0957 PT (worldtree-dev GO), restore drill passed.** The copies:
|
||||
- corviduo-dev `/var/lib/wt-legacy-archive/` (root 0700, UNENCRYPTED): demo 51 files and personal 1,426 files, each as tar.zst plus per-file sha256 plus MANIFEST.txt.
|
||||
- A DEDICATED restic repo, rest-server-nh3 `/nh3-dev/wt-legacy-archive/`, snapshot `98dc64e0`, with its password vaulted at `nh3-dev/wt-legacy-archive/restic-password`. Its URL is the vaulted `nh3-dev/etc/restic/repository` plus `wt-legacy-archive/`. It is mirrored to ana-nas at 05:00.
|
||||
- It is deliberately NOT in the main nh3-dev /home sweep, whose keep-monthly 12 retention would break the contract's 30-day rule.
|
||||
- The drill: 51/51 and 1,426/1,426 per-file sha OK, sqlite integrity ok, and a one-flipped-byte negative control was caught.
|
||||
- **Contract rev 1.2: DESTROY WHOLE at retirement-done or 2026-10-30, whichever is first, or on any subject-erasure request.** The runbook (sent to worldtree-dev):
|
||||
1. rm `/var/lib/wt-legacy-archive` on corviduo-dev.
|
||||
2. rm `/volume1/Backup/restic/nh3-dev/wt-legacy-archive` on nh3-nas.
|
||||
3. The ana-nas mirror's `--delete` removes that copy; verify it.
|
||||
4. `secret rm` the password AND purge Vaultwarden's trash (crypto-shred), and check for NAS share snapshots.
|
||||
- Observation: the context_promotion dir's mtime moves at boot even under mode=off, while its files do not; this was flagged to worldtree-dev for C3.
|
||||
- **The task:** replace the live speech seat with `nvidia/parakeet-unified-en-0.6b` under NeMo 3.0.0 with bf16 weights. The **NVIDIA Open Model License is ACCEPTED** for internal use.
|
||||
- The live seat today: container `parakeet` on fv-ml1 GPU 0, port :8300, sherpa-onnx int8 `parakeet-tdt-0.6b-v3`, reached through LiteLLM as `ext-stt` and `whisper-1`; its caller is `talk`.
|
||||
- Evidence, `docs/pfi/parakeet-seat-ab-2026-09-30.md` (a6c1d3c): end-to-end p50 for 1–3 / 3–8 / 8–20 s clips goes 144 / 260 / 565 → 23 / 27 / 33 ms, and WER is lower on every set.
|
||||
- **Kit:** the wrapper `services/parakeet-ab-2026-09-30/code/serve_nemo.py` keeps the seat's endpoints and text, and matched NeMo's own transcribe on 400/400. The weights are pinned on fv-ml1 in `/tank/aimodels/huggingface` (rev `fe53cd88`). A working NeMo 3.0.0 env for reference is under `/tank/spikes/parakeet-ab`. **No image is built yet.**
|
||||
- **What the image needs:**
|
||||
- a warm-up at the longest served length;
|
||||
- a bf16 cast BEFORE `.to(cuda)`, which avoids a +1.5 GB load spike;
|
||||
- local attention for long files (a 30-min file took 2.6 s in one request).
|
||||
- **Room:** it needs about +1.1 GB while serving (+1.5 GB at load) over the seat's 1,690 MiB, and GPU 0 has ~100 MiB free.
|
||||
- The plan is to trim `vllm-gen-small` `--gpu-memory-utilization` from 0.48 to about 0.46 at a quiet moment (a 2–3 min restart).
|
||||
- ⚠ The util value does NOT predict resident VRAM: on 09-15 gen-small at 0.48 held 36,942 MiB and cyberprev at 0.40 held 47,124. MEASURE nvidia-smi Free after the change; do not compute it. My 09-30 "0.01 ≈ 0.95 GB" estimate is unverified.
|
||||
- GPU 1's ~6.6 GB free is intern-decision's 32k headroom, so it is not available.
|
||||
- **Cut-over:** keep the old seat as the rollback, and leave LiteLLM alone unless the port changes. Re-measure live on GPU 0: latency per length bin against the old seat, a WER spot-check, and memory.
|
||||
- **Live-seat defects until then:**
|
||||
- HTTP 500 above ~400 s of audio;
|
||||
- long-form dropouts;
|
||||
- after a 1.5 s digital-silence pause it can drop the rest of the utterance (6 of 40).
|
||||
|
||||
### Worldtree U11 memory cutover (demo + personal)
|
||||
|
||||
- **Legacy plane OFF since 0115/0120 PT 2026-09-30** (config repo 63cf268; personal /metrics in b6fdd81). → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
|
||||
- **Daily gate batches:** infra-hermes runs them from 2026-10-01 with `scripts/wt-memory-gate-batch` and copies me on every verdict. The count is **1 of 3** consecutive PASS at off (20260930T090608Z); a FAIL restarts it.
|
||||
- **⚠ U11b STEP 5 IS MINE, triggered by the 3rd consecutive PASS:**
|
||||
1. Run `docker exec -i <api> python - < scripts/wt-h2-count.py` VERBATIM, right before each instance's deletion. Exit 2 means STOP and send worldtree-dev the output.
|
||||
2. Delete LIVE, with the api running, using literal paths only: `agents/{forseti,lofn,mimir}/memory/<agent>.chroma` and `memory/context_promotion`, on BOTH instances.
|
||||
3. Send worldtree-dev the stamp; b193 ships after it.
|
||||
- A b192 restart re-creates an empty schema-only `ledger.db`. That is residue, not memory data: say so in the stamp, and remove it after b193.
|
||||
- **Legacy archive:** DESTROY it whole by 2026-10-30, or at retirement-done, or on any subject-erasure request, whichever comes first. The runbook is in the detail file.
|
||||
- **TODO:** re-sweep both api logs after real traffic, grepping `legacy executor|Traceback|ERROR|will not be remembered`. After the b193 push, remove the retired config keys at my pace.
|
||||
|
||||
### fv-ml1 GPU layout (as of 2026-09-30)
|
||||
|
||||
- **GPU 0:** cyberprev (47.1 GB), gen-small (37.5 GB), voices (10.8 GB), the parakeet seat (1.7 GB); ~100 MiB free.
|
||||
- **GPU 1:** vllm-coder, erp-seat, meromero-rp, plus intern-decision (cap 14.4 GiB, 32k tokens, peak 15,220 of a 15,437 MiB budget). FULL.
|
||||
- **GPU 3:** the full-size-seat reserve (Flash-Next is parked). On-demand tenants: Blender, and Scriberr (0 idle, ~5.5 GB per job). When a full-size seat claims GPU 3, Scriberr steps aside to **irv-ml1's A6000**, not back to GPU 1.
|
||||
|
||||
### intern-decision (replaced SemIf on 2026-09-30)
|
||||
|
||||
- **LIVE 0.1.3** at `intern-decision.fv.internal:8033`: semif-compatible `/decide` plus Jev `/v1/systemone`, 32k tokens, a Triton cache volume. Run `scripts/intern-decision-warmup` after an IMAGE change. → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
|
||||
- **Open:** label ~50 real Wyrd/Cicada turns before trusting it in production; its card makes no contamination claim.
|
||||
|
||||
### Scriberr (fv-ml1 GPU 3)
|
||||
|
||||
- **LIVE `scriberr:local-blackwell-a353078-dropout2`:** upstream a353078 plus patch 0001 (overlap slicer) and patch 0002 (gap retry, `PARAKEET_MODEL_PATH`). v3 stays (Prime: no NeMo 3.0.0 surgery). → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
|
||||
- **Awaiting Prime:**
|
||||
- Open the slicer upstream PR, and choose which GitHub account (`stacks/scriberr/patches/upstream-pr/`).
|
||||
- Delete the leftovers:
|
||||
- the candidate weights in `/tank/aimodels/huggingface`, EXCEPT unified-en, which the seat switch needs;
|
||||
- `/tank/spikes/scriberr-slicer`, including `private/`, which holds Prime's recordings (mode 700);
|
||||
- `/tank/spikes/parakeet-ab` (~25 GB), but only after the switch.
|
||||
|
||||
### irv-ml1 /storetank at 86% (2026-09-30)
|
||||
|
||||
- infra-hermes and comfy-dev built a 4-tier reclaim plan (thread `01M3TCSYRSFNAPA9BPQTMFQ6KJ`):
|
||||
- A, staging: 29 GB;
|
||||
- B, identical duplicates: ~9+ GB;
|
||||
- C, superseded generations: ~80 GB;
|
||||
- D, the June wave: ~45 GB.
|
||||
- Delete-hold is ON, awaiting Prime. I recommended A and B. 261 GB free.
|
||||
|
||||
### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
|
||||
|
||||
@@ -176,66 +215,6 @@ _As of 2026-09-30 ~0120 PT._
|
||||
- worldtree-instance-configs: its 6 long-unpushed commits (53349f8…6d4ac44, Aug 2 to Sep 27) were
|
||||
**pushed on Prime's go-ahead** (a9d091e..6d4ac44). Origin, the repo and both hosts now agree.
|
||||
|
||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||
|
||||
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
||||
- **Scriberr slicer patch LIVE 2026-09-30 1211 PT** as `scriberr:local-blackwell-a353078-slicer1` (Prime: "build the slicer"). Chunks now overlap by 4 s inside the 120 s and hand over at a word both transcribed; that took cuts with an error nearby from 52 % to 22 % against a 19 % background (floor ±0.08, 4 files × 3 placements). Pause-aware cutting measured neutral, so it is opt-in (`--pause-search`). The brief's start-time stitch duplicated words at a quarter of the stitches, which is why the handover is by agreed word. Peak 5,496 MiB (GPU 3 n=3, live GPU 1 n=1). Rollback: `SCRIBERR_IMAGE=scriberr:local-blackwell`, `.env.bak-20260930-pre-slicer1`. Upgrade: `scripts/scriberr-rebuild --sha <sha>`. Contract: `stacks/scriberr/patches/README.md`; bench: `docs/pfi/scriberr-slicer-bench-2026-09-30.md`.
|
||||
- **Upstream PR prepared, NOT opened; it needs Prime's yes** (`stacks/scriberr/patches/upstream-pr/PR.md`).
|
||||
- **Dropout INVESTIGATED 2026-09-30 (Prime via coordinator; investigation only, nothing deployed):** `docs/pfi/parakeet-dropout-investigation-2026-09-30.md`. Real losses against ground truth (SCOTUS official transcript, Gutenberg #38916): v3 loses ~140 / 66 / 50 / 51 clean words per transcript (audiobook / argument / p1 / p2). Cause = v2/v3 0.6B weights collapse deep in long full-attention windows (encoder-side; 1.1B TDT/RNNT/CTC never do). No decoding, context, loudness or resampling fix. **Fix = re-transcribe ≥3 s speech gaps: −80–90 % everywhere** → `stacks/scriberr/patches/proposed/0002` (+ `PARAKEET_MODEL_PATH`), built as `scriberr:local-blackwell-a353078-dropout2`, NOT deployed; peak 5,506 MiB. Prime's calls: ship 0002?; v2 (0 on his files, collapses on read speech) vs keep v3; parakeet-unified-en-0.6b (needs NeMo 3.0.0 + NVIDIA Open Model License). Weights pulled pinned into `/tank/aimodels/huggingface` (~30 GB); throwaway env `/tank/spikes/scriberr-slicer/envs/nemo300`.
|
||||
- Scriberr moved to **fv-ml1 GPU 3** (coordinator, 2026-09-30); `scriberr-rebuild` memory stage now counts only its own PIDs and needs ≥20 GB free. Its default budget is still the retired GPU 1 5,496 MiB (0002 peaks 5,506 → pass `--budget`).
|
||||
- Private bench data (copies of Prime's two uploads + transcripts) sits in fv-ml1 `/tank/spikes/scriberr-slicer/private/` (mode 700), kept pending Prime; the public audio and metrics are beside it.
|
||||
- **Parakeet SEAT A/B DONE 1745 2026-09-30** (`docs/pfi/parakeet-seat-ab-2026-09-30.md`, a6c1d3c). **The seat's latency is its RUNTIME, not its model:** the sherpa-onnx int8 graph runs on ONE CPU thread with the GPU at 2–9%.
|
||||
- End-to-end p50 for 1–3 / 3–8 / 8–20 s clips: seat 144 / 260 / 565 ms; `parakeet-unified-en-0.6b` under NeMo with bf16 weights 23 / 27 / 33 ms. The floor is ≤ 6 ms, and a +50 ms positive control read +52.
|
||||
- Unified also wins English WER everywhere: LS-clean 1.97 against 2.70, LS-other 3.09 against 4.56, AMI 8.30 against 12.69.
|
||||
- Unified int8 in the seat's runtime is SLOWER, so the runtime has to change.
|
||||
- **Live-seat defects:** HTTP 500 above ~400 s (the ONNX position table is fixed at 5,000 frames); long-form dropouts of 320–1,676 of 3,580 words on 6-minute files; and after a 1.5 s digital-silence pause it can drop the rest of the utterance (6 of 40; `talk` is a caller).
|
||||
- **Fit:** unified bf16 needs +1.1 GB serving (+1.5 at load) over the seat's 1,690 on GPU 0. ⚠ GPU 1's 6,625 free is NOT spare: it is intern-decision's 32k headroom.
|
||||
- Switch kit: NeMo 3.0.0 plus `services/parakeet-ab-2026-09-30/code/serve_nemo.py` (same endpoints and text); the image is NOT built; it needs a warm-up, a bf16 cast before moving to the GPU, and local attention for long files. Licence: NVIDIA Open Model License. The spike dir is fv-ml1 `/tank/spikes/parakeet-ab` (~25 GB).
|
||||
- **Awaiting Prime:** the switch, the gen-small KV trim (~2 GB, util 0.48→0.46), the licence, and a `NUM_THREADS=16` stopgap (~40% faster).
|
||||
- **Scriberr gap-retry fix (patch 0002) LIVE 1602 2026-09-30** as `scriberr:local-blackwell-a353078-dropout2` (Prime: "basic fix, no surgery for the new toolkit"; v3 stays). The investigation (`docs/pfi/parakeet-dropout-investigation-2026-09-30.md`) found the v3 drops are real against ground truth (50–140 words per transcript); the retry cuts them 80–90%. Live check: 5,502 MiB, `retried_gaps` reported. Rollback: `.env.bak-20260930-pre-dropout2` / slicer1. Leftovers kept pending Prime: 26 GB of candidate weights, `envs/nemo300`, and the private bench data in `/tank/spikes/scriberr-slicer/`.
|
||||
- **2026-09-30 1322–1335, Prime: "Go GPU 3 now and extend the jev endpoint to hit 32k tokens".** DONE.
|
||||
- **Scriberr is on fv-ml1 GPU 3** (`SCRIBERR_GPU_ID=3`; a 20-min file verified at 5,496 MiB). It is an on-demand tenant of the reserve, like Blender: it STEPS ASIDE when a full-size seat claims GPU 3, and it goes to **irv-ml1's A6000**, NOT back to GPU 1.
|
||||
- **intern-decision: `VRAM_CAP_GIB=14.4`, `MAX_TOKENS=32768`** (Jev's 32k). The measured card peak at the limit is 15,220 MiB (1 and 16 questions, n=3) against a 15,437 budget; 32,769 tokens → 422; latency 2.1 s at 32k. JevBench is still 202/231 with 0 diffs.
|
||||
- ⚠ The first call in a new length bucket costs ~6.5 s: Triton/fla autotune, apparently in 2,048-token buckets, ~16 of them up to 32k. The cache (`/tmp/triton-cache`, writable layer) SURVIVES `docker restart` (measured) but is LOST on recreate. **DONE as 0.1.3 (infra-hermes f65e27b), and my AUDIT PASSED 1604:** the named volume `intern-decision_triton-cache` survives a force-recreate (6 random sizes from 5.5k to 32.6k all warm), and the 16 × 2,048-token buckets are PROVEN. Run `scripts/intern-decision-warmup` after an IMAGE change only; cold it takes 109 s, warm 17 s. JevBench still has 0 diffs.
|
||||
- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).**
|
||||
- Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`.
|
||||
- Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16.
|
||||
- Limits: `VRAM_CAP_GIB=9.0` with `MAX_TOKENS=7168`, which returns 422 up front. Rest 8.8 GB, card peak 9,866 MiB. Latency 80 ms server-side for 21 criteria and 205 ms for 16 × 3.9k.
|
||||
- Acceptance on the live URL: 240/259 pooled and 79/84 Wyrd, with 0 of 560 rows changed against the bench.
|
||||
- The semif container was REMOVED at 0949 via `compose down`; the image, files and token are kept (rollback in `stacks/semif/README.md`). There were no semif consumers to migrate.
|
||||
- Still open: label ~50 real Wyrd/Cicada turns before trusting it in production (the card makes no contamination claim).
|
||||
- **Jev API: the MODEL speaks it, the SERVICE does not.** Jev is TypeSafe's closed `jev-latest`: `POST /v1/systemone {state, model, questions:{id:{type noul|choice|score, instructions, criteria}}}` → `{answers:{id:{type, noul | choice+probabilities | probabilities}}, usage, model}`. The bench drove Intern-Decision natively through JevBench's `typesafe` adapter, 231 items × 4 repeats, all OK, so the engine's I/O matches that subset. intern-decision-serve exposes ONLY semif's `/decide` and `/decide/shared`, so a Jev client gets a 404. Adding `/v1/systemone` is a thin passthrough (the bench's 60-line wrapper is the seed); images would still be refused, because the vision tower is dropped for the GPU budget. **LIVE as 0.1.1 (infra-hermes ff552ab, deployed 1255 PT); infra-ops AUDIT PASSED at 1310.** My independent JevBench v1.2.16 typesafe run against the live endpoint: 202/231, hard 83/111, 0 row diffs against bench r1..r4 (the positive control, native vs drop-in, shows 11 diffs). Two low findings went back to infra-hermes: the contract's example shows `model` as an object while the wire uses a string, and `images: []` is wrongly refused. **Fixed in 0.1.2 (1866c00, live about 1303 PT), and my re-audit PASSED:** `images` [] and null are treated as absent, non-empty is 422, JevBench is still 202/231 with 0 diffs, and 124 tests pass. The rollback chain is 0.1.1 → 0.1.0. **True Jev, per docs.typesafe.ai/models + /api (read 2026-09-30): TEXT ONLY** ("No image, audio, or video input"), so our images→422 matches it. Jev 1.13 allows 64k tokens per request (32k for state plus the longest question) and up to 255 options per choice; its errors are 422, 429 and 529. **Our real gap is context: MAX_TOKENS 7,168 against Jev's 64k,** set by the GPU 1 memory cap; a Jev client with a big state gets 422. The rest was probed live and matches: `usage` has input_tokens/output_tokens, instructions accepts an object or an array, and a 20-option choice returns 200. The pass line: JevBench v1.2.16's `typesafe` adapter against the live endpoint gives 202/231 (hard 83), with 0 row diffs against the bench ledgers. Also: >16 questions → 422 (no chunking), images → 422, peak ≤ 9,876 MiB, one inference thread only. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback.
|
||||
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. The losing candidate weights (plumb-4b, JevK5 v0.2+v0.3, imajev-4b; about 24 GB) were DELETED on Prime's word at 1234 2026-09-30; Intern-Decision-4B is kept because it is live, and the pinned SHAs for a re-pull are in the bench doc. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
|
||||
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||
contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1%
|
||||
(95% CI +5.1..+14.3); unanimous rotations 94.5% accurate. Latency, envelope and the fast-kernel A/B
|
||||
are in `stacks/semif/README.md`.
|
||||
- The heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`) is triaged and folded into 0.1.3:
|
||||
C1–C6, S1–S3, S5, S6, S8–S10 fixed with tests; S4 settled; S7 (publish on all interfaces) accepted
|
||||
as LAN + auth.
|
||||
- NVFP4 is not worth it (see the SemIf detail files). Prime's probes (2AM/2PM, dragon/lottery) are
|
||||
recorded in [[2026-09-27-semif-order-averaging]].
|
||||
- **Consumer-fit spikes DONE (Prime, 0904 — scope was Wyrd scene change + Cicada emotion):** nothing
|
||||
built, two calls are Prime's. Cicada: an input-only "does this earn a reaction?" gate scored 30/31
|
||||
with descriptive options and 19/31 with terse yes/no, so the wording carries it. Proposed as a
|
||||
gesture-only gate in talk `/face`. It overrides the model's affect, which Cicada's 2026-09-20
|
||||
ruling reserves. Wyrd: fits the CHOICE, not the writing. A location-anchored "left this place?"
|
||||
gate scored 21/21 after the first wording failed its controls; exit choice scored 18/21. Parked
|
||||
unless live play shows node churn. → [[2026-09-27-semif-consumer-fit-spikes]]
|
||||
- **Rulings (Prime, 0937): both spikes, build nothing.** Follow-up on SemIf as Cicada's WHOLE
|
||||
mood source (henge id 88): **not faster and does not work as well.** First paragraph +32 ms async
|
||||
and +94 ms sequential vs today's 246 ms, because the pose header costs only ~31 ms and SemIf
|
||||
shares GPU 1 with the LLM. Acceptable pose 67% vs 92%; the mood carried 7/15 vs 14/15. Upside:
|
||||
gestures at 13% vs 58%.
|
||||
- **Prime 0948: idea 88 dropped; fix the 422 → DONE, 0.1.4 live (the `fix(semif): 0.1.4` commit).** INV-7 wraps SemIf's
|
||||
`shared._state_prefix` so the prefix is only the tokens the full prompts share. Startup proves the fix
|
||||
is in effect (the hook must be what score_shared resolves, the prefix unchanged on an ordinary state,
|
||||
and a merge-prone state scored through the shared path). Folded from heid bug hunt SKAL (Hulda, thread
|
||||
`01M3HXMXN27F3K534Q6QS45AHV`). Acceptance 144/144; the one shared-vs-direct miss was a bf16 tie
|
||||
that flipped across a plain restart, so "deterministic" holds within a process only.
|
||||
|
||||
### Blender on fv-ml1 GPU 3, agent-driven (2026-09-27, Prime)
|
||||
|
||||
- Prime: "go ahead with gpu 3, both", and he does not use Blender, so **agents drive it through MCP**.
|
||||
@@ -321,14 +300,18 @@ _As of 2026-09-30 ~0120 PT._
|
||||
|
||||
### Live threads
|
||||
|
||||
- git: `main` == origin at `274b817` (pushed 2026-09-29 on Prime's word), plus this snapshot. `graphify-out/GRAPH_REPORT.md` stays modified
|
||||
and uncommitted on purpose: it is auto-regenerated.
|
||||
- git: origin/main is at `128d1d8`, pushed 2026-09-30 1047 by someone other than infra-ops (presumably Prime). Local is ahead with unpushed commits, this snapshot included. `worldtree-instance-configs` has 3 unpushed commits (0a1387e, 63cf268, b6fdd81). Pushing is Prime's call. ⚠ The working tree AND index are shared with infra-hermes and subagents: commit with `git commit -- <paths>` (auto-memory `feedback_shared_git_index_commit_pathspecs`). `graphify-out/GRAPH_REPORT.md` stays modified and uncommitted on purpose: it is auto-regenerated.
|
||||
- nh3-dev root disk was cleaned 2026-09-30 1704 (uv prune, dangling images, old build cache): 86% → 82%. The Beszel 85% alert flaps near the line.
|
||||
- Booth submit-all fix (Prime's report) is LIVE since 2026-09-27 1705, via booth-dev (booth `50bfc7b`).
|
||||
Pushing it is booth's call, per Prime; it is not ours.
|
||||
- ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-30]` **Parakeet speech seat → `parakeet-unified-en-0.6b` under NeMo (bf16) APPROVED by Prime, NVIDIA Open Model License accepted. Implementation is deferred to the next session, tracked by the in-flight "NEXT" section and a6c1d3c.** → `persistent-memory.d/2026-09-30-parakeet-seat-switch-approved.md`
|
||||
- `[2026-09-30]` **Worldtree U11a: legacy memory plane OFF on demo and personal. The U11b data deletion is gated on 3 consecutive PASS and step 5 is mine; the legacy archive must be destroyed by 2026-10-30.** → `persistent-memory.d/2026-09-30-worldtree-u11a-off-u11b-gate.md`
|
||||
- `[2026-09-30]` **SemIf replaced by intern-decision (Intern-Decision-4B, the Jev bench pick): semif-compatible plus Jev `/v1/systemone` at 32k tokens on GPU 1, with a Triton warm-up cache volume.** → `persistent-memory.d/2026-09-30-semif-replaced-by-intern-decision.md`
|
||||
- `[2026-09-30]` **Scriberr moved to GPU 3 (on demand); our build carries the overlap slicer (0001) and the Parakeet gap retry (0002); v3 kept.** → `persistent-memory.d/2026-09-30-scriberr-slicer-gap-retry-gpu3.md`
|
||||
- `[2026-09-29]` **Worldtree U11a prepped, not flipped: a staged demo config plus the agreed U8 window plan (infra-hermes runs the batches).** → `persistent-memory.d/2026-09-29-worldtree-u11a-prepped.md`
|
||||
- `[2026-09-28]` **Worldtree U10 backfill done on demo (5) and personal (797). model_roles drift needed a memory_tagger sync first; mimir had missing vectors.** → `persistent-memory.d/2026-09-28-worldtree-u10-backfill.md`
|
||||
- `[2026-09-28]` **Bonsai ternary vs Q4_K_XL at concurrency on the 275 W card (1.93x at N=1 falls to 1.06x at N=8, 1.21x with the MMVQ fix); the weights are acquired.** → `persistent-memory.d/2026-09-28-bonsai-ternary-spike.md`
|
||||
@@ -487,67 +470,27 @@ _As of 2026-09-30 ~0120 PT._
|
||||
- `[2026-09-16]` ⭐ **voices-seat live: one carrier, N `lv-<author>` LoRA adapters, hot-swap measured at 0.24 s.** LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying; `--gpu-memory-utilization` is a request against TOTAL VRAM and only a pinned KV makes it predictive. → `persistent-memory.d/2026-09-16-voices-seat-lora.md`
|
||||
- `[2026-09-16]` ⭐ **lv-hemingway corpus gated at 994,760 words — and half the catalogue had to be EXCLUDED.** 169,759 words of measured 90–96% collection duplication, a Sherwood Anderson parody, and the author's own name 95 times in publisher back matter; the gender resolver needed a corpus base-rate correction to stop reading women as men. → `persistent-memory.d/2026-09-16-lv-hemingway-corpus.md`
|
||||
- `[2026-09-16]` **Grok token broker built then SHELVED — operator ruled "keep the jail stop the a/b", so the renewal feature has no consumer.** ⛔ Do NOT arm `probe-rotation`: the risk did not shrink (it reaches BOTH Gróa transports through one shared session) and the payoff went to zero. → `persistent-memory.d/2026-09-16-grok-broker-shelved.md`
|
||||
- `[2026-09-15]` ⚠⚠ **DO NOT carry "a client-side timeout is not a cancellation" as a rule — it is FALSE as stated.** A clean abandon cancels itself ~6 s later (measured); yet six requests genuinely orphaned on `vllm-erp-seat`. Some propagate, some do not, **boundary unknown** — which argues for a detector, not a rule. ⭐⭐ The durable artifact: **a serving engine's KV cache CYCLES, an orphaned one only CLIMBS** — request count and throughput are ambiguous between loaded and wedged, and I called the seat healthy twice off them (correctly, on the evidence). ⚠ A `max_tokens` ceiling would NOT have prevented it: the worst offender had 16384 set, hit it, and returned 24,594 chars of whitespace. → `persistent-memory.d/2026-09-15-client-abandon-cancellation-boundary.md`
|
||||
- `[2026-09-15]` ⚠⚠ **`--gpu-memory-utilization` DOES NOT PREDICT RESIDENT VRAM — measure it, never compute it.** Wrong in **both** directions on fv-ml1: `vllm-cyberprev` util 0.40 (expect ~39,155 MiB) holds **47,124** (+8 GB over); `vllm-gen-small` util 0.48 (expect ~46,986) holds **36,942** (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Read `nvidia-smi --query-compute-apps`. Full per-seat residency table + the breeze shuffle arithmetic → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
|
||||
|
||||
- `[2026-09-15]` **breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75, `move-the-tts-stack-breeze-tts-bragi-tts-gateway`), triggered on evacuating embed/rerank/reward.** ⚠ Trigger as stated says "gpu0" but those three are on **GPU 1** (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because only `breeze-tts` is GPU-resident (~10.3 GiB, **growing**) while `bragi` and `tts-gateway` are CPU proxies, and co-location is what avoids a cross-site hop per TTS call. **breeze-tts sizing — original recommendation NOT to move it.** ~**10.3 GiB** measured under load at 53 min uptime, **up from 9.2 GiB** shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a **1.7 GB margin and shrinking**, on the live chat serving path. ⚠ Two measurement traps: it reports **nothing at idle on the wrong card** (`BREEZE_GPU_DEVICES=0` = the **3090**, not the A6000), and an early reading understates it. ⭐ The real objection is **topology**: `tts-gateway` is on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. → `persistent-memory.d/2026-09-15-breeze-placement-sizing.md`
|
||||
|
||||
- `[2026-09-15]` **Parakeet STT live on fv-ml1 GPU 0, behind LiteLLM `ext-stt` / `whisper-1`.** ⚠ **Placed on GPU 3 first, which was wrong — operator caught it.** A ~800 MiB seat should ride the card with the most uncommitted headroom (GPU 0, util 0.88, ~13 GB spare), not put the first fingerprint on the one pristine 96 GB card: vLLM sizes KV cache against TOTAL VRAM, so any tenant on an empty card eats a future full-size seat's profiling margin (flash-next needs 93 of 96 GiB). **GPU 3 is now a deliberate reserve at 2 MiB.** Retargeted the existing `stacks/parakeet/` (sherpa-onnx + our own FastAPI wrapper) from irv-ml1; v3 int8, 25 languages. ⚠ **ORT's CUDA EP compiles kernels lazily and the first decode on sm_120 took 45.7 s** — every later call ~0.5 s; a startup warmup in `app.py` now absorbs it, so the first real request is 0.65 s instead of a 45 s hang that no client would wait through. GPU use was **verified by a process on GPU 3 (922 MiB), not by the `provider=cuda` log line**, because ORT falls back to CPU silently and still returns correct text. Silence → `""` (null control), known sentence → near-exact (positive control). → `persistent-memory.d/2026-09-15-parakeet-stt-fv-ml1.md`
|
||||
|
||||
- `[2026-09-15]` ⭐⭐⭐ **THE FLEET'S CHARACTERISTIC FAILURE, named: a confident answer from a broken instrument.** Nine instances in one night, every one of which PASSED A CHECK — `provider=cuda` while ORT ran on CPU; `node --check` green on a file whose SERVED script was dead; `secret get` returning `""` with exit 0; `find()` turning a failed listing into an authoritative "not found"; a 401 rendering as "0 toolsets"; `compat` ✓ on a typo'd path; `doctor` exit 0 on ERROR; `ss | grep python` missing a listener named `hermes`; SIGTERM freeing a port 35 s before the process died. ⚠ **The tell: whenever "broken" and "legitimately empty/absent/off" produce the same output.** Remedies: measure the output not the input, positive AND true-negative controls, refuse to emit the ambiguous value, and never declare victory on a plausible fix. → `persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md`
|
||||
|
||||
- `[2026-09-15]` **`secret get` returned EMPTY with exit 0 under concurrency** — (svos-dev found it; 0/4 succeeded here). → `persistent-memory.d/2026-09-15-secret-get-returned-empty-with-exit-0-under-concurrency.md`
|
||||
|
||||
- `[2026-09-15]` ⭐⭐ **A check that reads an artifact AS STORED cannot see a transformation between storage and execution** — named twice in one night and it generalises. `node --check` on a source file passes while the SERVED page's inline script is dead (a JS `'didn\'t'` inside a Python string arrives as `'didn't'` and closes it); `provider=cuda` in a log echoes configured intent while ORT silently ran on CPU. Both check the INPUT to a transformation and get reported as checks of its OUTPUT. Remedy: gate the wire, not the file — `tts-stack tools/gate_served_page.py`. ⚠ My first version had a gap tts-dev closed: **a worklet inside a template literal is just a string to a parse of the enclosing script**, so its syntax error surfaces as a rejected `addModule` promise and *silent degradation*. I checked the instance, not the class. → `persistent-memory.d/2026-09-15-talk-v10-deploy.md`
|
||||
|
||||
- `[2026-09-15]` ⚠⚠ **The talk-deploy "permission problem" NEVER EXISTED — and I built a fix for it anyway.** `/opt/docker/compose` on nh3-dev is `root:docker 2775`, sessions run as `lkraven`, `lkraven` is in `docker`; a `mkdir` settles it in one second and nobody ran one for nine days. There is no `tts-dev` OS account at all. It held because a **stale memory row** supplied a mechanism, the operator's **routing instruction** ("give it to infra") was misread as corroboration of a *capability limit* — different claims, only one ever stated — and I **repeated it to the operator as fact**. Then, told to fix "the harness issue", I inferred an auto-mode classifier refusal and **committed a settings.json to tts-dev's repo on that inference**; their `mkdir` disproved it and I reverted. ⭐ **"I can't do X" is a hypothesis until someone pastes the error.** ⚠ That commit also overclaimed a doc fix that failed — **never chain an edit and its commit in one invocation.** → `persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md`
|
||||
|
||||
- `[2026-09-15]` **talk v10 LIVE on nh3-dev :8092 — the fleet speaks and listens on one page.** — First consumer of the `ext-stt` Parakeet seat: `POST /api/listen`, push-to-talk, barge-in. → `persistent-memory.d/2026-09-15-talk-v10-live-on-nh3-dev-8092-the-fleet-speaks-and-listens.md`
|
||||
|
||||
- `[2026-09-15]` **Two restart patterns from svos-dev worth stealing: (a) DRY-RUN BOOT against the still-held port** — start the new process while the old one holds the socket; it proves every check above the bind and dies on… → `persistent-memory.d/2026-09-15-two-restart-patterns-from-svos-dev-worth-stealing-a-dry-run.md`
|
||||
|
||||
- `[2026-09-15]` ⭐ **`svos_miranda` ENABLED and LIVE in Hermes — but `agent.disabled_toolsets` is permanently OFF by operator ruling ("i dont want the tools disabled everywhere").** That key is a **global** end-of-pipeline subtraction, not api_server-scoped: measured 46 tools → 20 on a default session. It is also **unnecessary** — `platform_toolsets.api_server: [svos_miranda]` alone resolves an api_server session to exactly the 8 tools, write-klass absent. Gateway restarted 02:10 (PID 3107822→3901622, observed); `/v1/toolsets` now 29 rows incl. `svos_miranda`; operator's own surface verified intact at 46. ⚠ **SVOS must stop verifying against the GLOBAL roster before it restarts** — it will see 29 and refuse, by design now. → `persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md`
|
||||
|
||||
- `[2026-09-15]` **irv-ml1 parakeet RETIRED; voice-studio STOPPED.** — Both operator rulings. → `persistent-memory.d/2026-09-15-irv-ml1-parakeet-retired-voice-studio-stopped.md`
|
||||
|
||||
- `[2026-09-15]` **`svos_miranda` Hermes plugin validated; found its load blocker.** Absolute intra-package imports (`from hermes_plugin.x`) could not resolve at the documented install name — fixed by svos-dev at `c964e64`. ⚠ **`hermes plugins validate` and `doctor` can NEVER pass this plugin**, by construction: validate's probe stub is config-blind AND returns `None` from `register_tool` (which the plugin's guard reads as a collision), and doctor runs under a temp `HERMES_HOME` with no config. ⚠ `doctor` exits **0** on ERROR (use `--ci`); `compat` reads a **nonexistent path as a pass**. Roster verified 8/7 by a probe supplying real settings. → `persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md`
|
||||
|
||||
- `[2026-09-15]` **ana-docker resolves NO `.internal` names** — its `/etc/resolv.conf` is `1.1.1.1`/`1.0.0.1`, not the fleet AdGuard. → `persistent-memory.d/2026-09-15-ana-docker-resolves-no-internal-names.md`
|
||||
|
||||
- `[2026-09-15]` ⚠⚠ **irv-ml1 still points at the retired wg0 lifeline `10.100.79.3` in 96 places — and one is a LIVE breakage, not a dead link.** `voice-studio` cannot reach `studio-gate` (both up, separate docker networks, gate URL is the dead IP) and has been failing since the 2026-09-06 cutover with nothing alerting. 8 running containers carry dead `homepage.href` labels; `waterland-studio`'s siteMonitor too. ✅ `tts-gateway`/`ext-tts` verified UNAFFECTED. Not fixed — wants a scheduled pass, not a 02:00 improvisation. ⭐ Third instance of the same shape: **a retired address needs a grep by ADDRESS, not by hostname, and labels live in no file until the container is recreated.** → `persistent-memory.d/2026-09-15-irv-ml1-dead-wg0-address.md`
|
||||
|
||||
- `[2026-09-15]` **Parakeet bench settled by tts-dev — FV wins at both clip lengths and beats the incumbent Whisper; IRV seat is now retirable.** → `persistent-memory.d/2026-09-15-parakeet-bench-settled-by-tts-dev-fv-wins-at-both-clip.md`
|
||||
|
||||
- `[2026-09-15]` **Mesh membership retired for fv-ml1 and nh3-dev — six nodes left, each with a job.** fv-ml1 gets break-glass rejoin instead of standing membership; nh3-dev's retirement also removed the nh3-scale masquerade exception it had required. Exactly one live reusable pre-auth key remains fleet-wide. → `persistent-memory.d/2026-09-15-fv-mesh-watchdog.md`
|
||||
|
||||
- `[2026-09-15]` **FV cross-site routing fixed — one OPNsense outbound-NAT rule had been scoped to Anaheim only.** fv-ml1 now reaches NH3/ESH/IRV/ANA/mesh/internet; four rules, all `src=10.251.50.0/24`. The diagnostic signature is the valuable part: every layer looks correct and the discriminator is that *every other site pair works*. → `persistent-memory.d/2026-09-15-fv-cross-site-snat.md`
|
||||
|
||||
- `[2026-09-15]` **Break-glass mesh path on fv-ml1** — inverted from a restore-watchdog on the operator's suggestion: the box is OFF the mesh and the watchdog JOINS it on fleet loss. Exposed a rejoin key expiring in 4 days; replaced with a dedicated 1-year key and the two stale reusable keys retired. → `persistent-memory.d/2026-09-15-fv-mesh-watchdog.md`
|
||||
|
||||
- `[2026-09-15]` **Fleet identity/group/path conventions pinned + docker trees → `root:docker 2775` setgid on 5 hosts.** `svc-*` in 800-849, infra-ops 850, docker 851, `vh` for new hosts with no retro-renames; `0777` cleared; `linus` deleted; `llmuser` de-privileged. → `persistent-memory.d/2026-09-15-fleet-identity-conventions.md`
|
||||
|
||||
- `[2026-09-15]` **nh3-dev unreachable from the mesh at its LAN address — Tailscale's `ts-input` anti-spoof, not DNS.** Fixed with a masquerade exception on nh3-scale. ⚠ Do NOT instead advertise the /32 from nh3-dev; that black-holes it from every other site while its own LAN keeps working. → `persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md`
|
||||
|
||||
- `[2026-09-15]` **ESPHome pinned to 2026.8.2 + `kb` KB-search tool shipped.** Untagged image had drifted a year; config relocated into restic with 539 MB of regenerable cache excluded; remote-build disabled (⚠ two switches, only one closes the port). `kb` exists because Worldtree's `/search` searches messages, not notes, and returns a clean empty result for a note that exists. → `persistent-memory.d/2026-09-15-esphome-and-kb.md`
|
||||
|
||||
- `[2026-09-15]` **Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos `7165272`)** → `persistent-memory.d/2026-09-15-hermes-bearer-rotation-hold-released-svos-dev-split-their.md`
|
||||
|
||||
- `[2026-09-13]` **STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint.** → `persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md`
|
||||
|
||||
- `[2026-09-11]` **Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m** → `persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md`
|
||||
|
||||
- `[2026-09-03]` **Run 3c is STAGED on pfi-gx10 and deliberately NOT launched** — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is *measured* inert. ⚠ The encode-cache FILENAME differs by design (`base_model_path` is in the key) — input hash, not output. ⚠ Tripped the `pkill -f` ssh self-match again; the launcher guards on a pidfile because of it. → `persistent-memory.d/2026-09-03-gx10-run3c-staged.md`
|
||||
|
||||
- `[2026-08-25]` **Fused MoE kernel path — DEFERRED, tracked at park `fused-moe-kernel-path-for-gemma-4-moe-training` (id 47).** → `persistent-memory.d/2026-08-25-fused-moe-kernel-path-deferred-tracked-at-park-fused-moe.md`
|
||||
|
||||
- `[2026-08-24]` **`nconnect=8` on `/mnt/smithy` — approved but DEFERRED at operator instruction.** brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread `01M0R46SFYF83099N16WD67KGD`.
|
||||
|
||||
- `[2026-08-19]` **AI-tab Dormant regrouping BELAYED by the operator** — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong… → `persistent-memory.d/2026-08-19-ai-tab-dormant-regrouping-belayed-by-the-operator.md`
|
||||
|
||||
_112 older entries archived to archival-memory.md._
|
||||
_135 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-09-30]` **Shorter Parakeet slices as Scriberr's memory fix.** I predicted ~6× less memory from the attention-matrix arithmetic. Measured, 300 → 120 s only went 9,384 → 6,510 MiB: a ~5.6 GB fixed floor dominates. `expandable_segments:True` was the real lever (5,496). Measure the process peak; never extrapolate it from one tensor.
|
||||
- `[2026-09-30]` **Whole-file local-attention Parakeet in Scriberr (context 255/255).** OOM past 16 GB on a 35-min file. Local attention inside chunks is also non-deterministic run to run.
|
||||
- `[2026-09-30]` **Start-time midpoint stitching of overlapped Parakeet chunks.** It duplicated a word at 26 of 108 stitches, because Parakeet timestamps a post-pause word anywhere inside the pause. Hand over at a word both chunks agree on instead.
|
||||
- `[2026-09-30]` **int8 ONNX (sherpa-onnx) as the low-latency Parakeet runtime.** The int8 graph runs on ONE CPU thread with the GPU at 2–9%. Unified-en int8 was slower than the seat; fp32 ONNX was 4–12× faster, and NeMo was fastest.
|
||||
- `[2026-09-30]` **GPU budgets computed as total − used.** nvidia-smi `Free` is ~640 MiB lower per card (driver reserve). Budget from `Free`.
|
||||
- `[2026-09-27]` **The Blender MCP server on nh3-dev, reaching the add-on socket over an SSH tunnel.** Every tool worked except `get_viewport_screenshot`, which had Blender write a file for the SERVER to read ("Screenshot file was not created"). The server now runs inside the Blender container over ssh+docker-exec stdio. That also means no port is published. → `stacks/blender/README.md`
|
||||
- `[2026-09-27]` **A TCP connect as the "is Blender ready" probe.** docker-proxy accepts on a published port before the app behind it listens, so it said "answering" while Blender was still loading. The probe now asks the add-on to `ping`. The same trap applies to any service behind a published port.
|
||||
- `[2026-09-27]` **`log.exception()` in a GPU failure path** — the record keeps `exc_info`, so any retaining handler (pytest's capture does) pins the traceback's frames and tensors. Log `traceback.format_exc()` text instead (semif-serve `engine._guard`).
|
||||
@@ -569,10 +512,5 @@ _112 older entries archived to archival-memory.md._
|
||||
- `[2026-09-18]` **Routing SearXNG's egress through a SOCKS5 proxy on esh-scale** — one day live, reverted. It fixed nothing, and the reason I gave for reverting it was itself wrong: the rollback was the counterfactual and it falsified my own published claim. Kept reverted on its own merits (no measurable gain, added a hard ESH dependency for all fleet search). → `persistent-memory.d/2026-09-18-searxng-esh-egress-reverted.md`
|
||||
|
||||
- `[2026-09-18]` **`api_key: !ENV SEARXNG_BRAVE_API_KEY` in searxng settings** — this build has NO `!ENV` YAML constructor, so the file was unparseable and the container crash-looped ten… → `persistent-memory.d/2026-09-18-apikey-env-searxngbraveapikey-in-searxng-settings.md`
|
||||
- `[2026-09-15]` ⚠⚠ **Probing OPNsense API endpoints by POSTing at them — one was `/api/core/system/reboot` and it took the FV site dark for 3.5 min.** Endpoints are ACTIONS; a 200 means it ran. The call I wanted was documented in this repo's own `docs/pfi/opnsense-api-reference.md`. → `persistent-memory.d/2026-09-15-opnsense-api-reboot.md`
|
||||
|
||||
- `[2026-09-15]` **Advertising `10.100.10.50/32` from nh3-dev** to make its LAN address mesh-reachable — black-holed it from ESH/ANA/FV/IRV while its own LAN and the internet kept working, so a one-host check passes cleanly. `lookup 52` at rule priority 5270 beats `main` at 32766. Fix belongs at the router. → `persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md`
|
||||
|
||||
- `[2026-09-15]` **Remote-site MASQUERADE rules on nh3-scale** for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing.
|
||||
|
||||
_115 older entries archived to archival-memory.md._
|
||||
_118 older entries archived to archival-memory.md._
|
||||
|
||||
Reference in New Issue
Block a user