diff --git a/persistent-memory.md b/persistent-memory.md index e8f3887..c586a3a 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -166,7 +166,8 @@ _As of 2026-09-30 ~0120 PT._ ### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime) -- **LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging +- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness. +- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.** - 0.1.3 acceptance: 144/144 parity with upstream; averaging through the service 78.6% → 88.1% diff --git a/servers/fv-ml1/README.md b/servers/fv-ml1/README.md index 2f912b0..d3c500a 100644 --- a/servers/fv-ml1/README.md +++ b/servers/fv-ml1/README.md @@ -182,7 +182,7 @@ embed/rerank/reward trio. GPUs are pinned per container via **GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):** > ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`, -> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (:8032, ~8.7 GB resting, +> `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`, plus **`semif`** (**OFFLINE since 2026-09-30, Prime; see `stacks/semif`**; :8032, ~8.7 GB resting when up, > hard-capped at 12 GiB; `stacks/semif`, since 2026-09-27). Read the host > (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran > here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.) diff --git a/stacks/semif/README.md b/stacks/semif/README.md index 1271674..8ad0c47 100644 --- a/stacks/semif/README.md +++ b/stacks/semif/README.md @@ -1,5 +1,15 @@ # semif +> ⚠ **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now").** It was +> stopped (`docker compose stop`, NOT removed) to give scriberr its GPU 1 headroom back: +> scriberr's Parakeet path cuts audio into 5-minute slices and needs over 6 GB, and with +> SemIf resident it had ~6.7 GB and hit CUDA OOM on a 35-minute file. Stopping it moved +> GPU 1 from 91,052 to 81,806 MiB used. The image, weights, config and token are all kept; +> `unless-stopped` keeps it down across a reboot. **Do not restart it without Prime's word.** +> To bring it back: `cd /opt/docker/compose/semif && docker compose start`, but first +> make sure scriberr has its room (the durable fix, shortening scriberr's slice length, is +> deferred by Prime to later). Everything below describes the service as it was deployed. + **SemIf option-logit decisions** on **fv-ml1 GPU 1**, the utility card beside `vllm-coder`, the erp/meromero seats and scriberr. Deployed 2026-09-27 at Prime's request. Hosting assessment: infra-hermes, 2026-09-25.