diff --git a/persistent-memory.md b/persistent-memory.md index 01447bf..0e6215a 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -179,7 +179,13 @@ _As of 2026-09-30 ~0120 PT._ ### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime) - **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness. -- **⚠ Prime ~0840 2026-09-30: "replace semif with intern-decision now", on GPU 1, with Scriberr fixed alongside (his pick).** A background agent is building `intern-decision-serve` (semif-compatible API, :8033, `stacks/intern-decision`, token `intern-decision/api-token`) under a hard cap: total footprint ≤ 10,300 MiB. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback. +- **intern-decision LIVE on fv-ml1 GPU 1 since 0941 2026-09-30, REPLACING SemIf (Prime: "replace semif with intern-decision now", with Scriberr fixed alongside).** + - Where: `http://intern-decision.fv.internal:8033`, image `intern-decision-serve:0.1.0`, token `intern-decision/api-token`. Code and contract are in `services/intern-decision-serve/`, the stack in `stacks/intern-decision`. + - Surface: semif-compatible `/decide`, `/decide/shared`, `/health`. It has 12 documented deltas; the main one is that the questions in one call share a prompt, in calls of at most 16. + - Limits: `VRAM_CAP_GIB=9.0` with `MAX_TOKENS=7168`, which returns 422 up front. Rest 8.8 GB, card peak 9,866 MiB. Latency 80 ms server-side for 21 criteria and 205 ms for 16 × 3.9k. + - Acceptance on the live URL: 240/259 pooled and 79/84 Wyrd, with 0 of 560 rows changed against the bench. + - The semif container was REMOVED at 0949 via `compose down`; the image, files and token are kept (rollback in `stacks/semif/README.md`). There were no semif consumers to migrate. + - Still open: label ~50 real Wyrd/Cicada turns before trusting it in production (the card makes no contamination claim). **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback. - **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. About 33 GB of candidate weights stay on fv-ml1 `/tank/aimodels/huggingface/hub` pending Prime. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`. - **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code + diff --git a/stacks/semif/README.md b/stacks/semif/README.md index 3ca7fe6..d8031e5 100644 --- a/stacks/semif/README.md +++ b/stacks/semif/README.md @@ -1,21 +1,21 @@ # semif -> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled at ~0510 PT: "replace semif +> ⚠ **REPLACED by intern-decision (Prime, 2026-09-30).** Prime ruled that morning: "replace semif > with intern-decision now". The replacement is `stacks/intern-decision` (Intern-Decision-4B, live since 0941 PT, > `http://10.251.50.54:8033`, `intern-decision.fv.internal`). It keeps this service's HTTP surface > (`/decide`, `/decide/shared`, `/health`), so callers only change the URL and the token > (`secret get intern-decision/api-token`). Its README lists every deliberate difference. > -> The `semif` container has been **stopped, not removed**, since 0135 PT that day, when it was -> taken offline to give scriberr its GPU 1 headroom back. Its image (`semif-serve:0.1.4`), -> weights, config and token are all kept, and `unless-stopped` keeps it down across a reboot. +> The `semif` container was stopped at 0135 PT that day, to give scriberr its GPU 1 headroom back. +> It was **removed** at 0949 PT (`docker compose down`) so that Homepage stops listing a dead +> tile. The image (`semif-serve:0.1.4`), the weights, this stack's files and the token are all kept. > Port 8032 and `semif.fv.internal` stay reserved for it. > > **Rollback, only on Prime's word.** > 1. Stop intern-decision first. The two services do not fit GPU 1 together next to scriberr. > semif held 9.2 GB at rest and peaked at 12.9 GB. > `cd /opt/docker/compose/intern-decision && docker compose stop`. -> 2. Start semif: `cd /opt/docker/compose/semif && docker compose start`. +> 2. Recreate semif: `cd /opt/docker/compose/semif && docker compose up -d` (the container is removed, so `start` will not work). > 3. Check nvidia-smi `Free` on GPU 1 against semif's 12.9 GB peak plus scriberr's 5.5 GB (its > 120 s slices) before calling it done. If it does not cover both, semif collides with scriberr. >