memory: U11b step-5 auto-trigger (3 PASS) is mine; semif->intern-decision in flight; scriberr GPU budget
This commit is contained in:
@@ -127,6 +127,7 @@ _As of 2026-09-30 ~0120 PT._
|
|||||||
- **PERSONAL flipped at 0120 PT** on 64f79b38a7fb, the same way. ⚠ Personal has **no /metrics**, because the #308 metrics block is demo-only, so the proof was the no-metrics one: an in-container `get_default→load_cutover_config→validate_cutover` gave off/writer/reader True, /health was 200, and the boot log matched the pre-flip one. Whether to add metrics to personal is still worldtree-dev's open call. Both instances are in sync with the repo.
|
- **PERSONAL flipped at 0120 PT** on 64f79b38a7fb, the same way. ⚠ Personal has **no /metrics**, because the #308 metrics block is demo-only, so the proof was the no-metrics one: an in-container `get_default→load_cutover_config→validate_cutover` gave off/writer/reader True, /health was 200, and the boot log matched the pre-flip one. Whether to add metrics to personal is still worldtree-dev's open call. Both instances are in sync with the repo.
|
||||||
- Watch after each flip: any `legacy executor <kind> invoked with memory.legacy.mode=off; ignored` line goes to worldtree-dev with its stamp. A sweep at 0154 of both logs since their last start found **0 watched lines, but there was ~no traffic** (demo 0 real requests, personal 1), so that proves little. **TODO: re-sweep after the first real traffic** with `docker logs --since <StartedAt>`, grepping `legacy executor|Traceback|ERROR|will not be remembered`. A CI recreate drops the old container's log.
|
- Watch after each flip: any `legacy executor <kind> invoked with memory.legacy.mode=off; ignored` line goes to worldtree-dev with its stamp. A sweep at 0154 of both logs since their last start found **0 watched lines, but there was ~no traffic** (demo 0 real requests, personal 1), so that proves little. **TODO: re-sweep after the first real traffic** with `docker logs --since <StartedAt>`, grepping `legacy executor|Traceback|ERROR|will not be remembered`. A CI recreate drops the old container's log.
|
||||||
- **Daily gate batches at off: LIVE, infra-hermes from 2026-10-01** (Prime's 2340 ruling; the plan was revised to `--legacy-mode off --label window-off` by worldtree-dev, thread `01M3RRW7WZ04GNR871HC1VKY4K`). The wrapper is `scripts/wt-memory-gate-batch`. It tests demo's deployed sha, from the detached worktree `~/development/Worldtree-gate` when worldtree-dev's tree has moved. Exit codes: 0 PASS, 1 FAIL, 2 error, 3 refused, 4 busy. **PASS and FAIL both go to worldtree-dev the same day**, because they build the n. **The U11b DATA deletion is gated on 3 consecutive PASS at off, or a Prime waiver**; code retirement is not gated. The controls were 080927Z FAIL and 082829Z FAIL (one flip each; the filing floor held; worldtree-dev calls it undecided). **Mine, 20260930T090608Z, PASSED (user median 0.83)**, so the count is 1 of 3 if the FAILs reset it. ⚠ It started by accident, from a test meant to be `--dry-run`, and matched the plan anyway.
|
- **Daily gate batches at off: LIVE, infra-hermes from 2026-10-01** (Prime's 2340 ruling; the plan was revised to `--legacy-mode off --label window-off` by worldtree-dev, thread `01M3RRW7WZ04GNR871HC1VKY4K`). The wrapper is `scripts/wt-memory-gate-batch`. It tests demo's deployed sha, from the detached worktree `~/development/Worldtree-gate` when worldtree-dev's tree has moved. Exit codes: 0 PASS, 1 FAIL, 2 error, 3 refused, 4 busy. **PASS and FAIL both go to worldtree-dev the same day**, because they build the n. **The U11b DATA deletion is gated on 3 consecutive PASS at off, or a Prime waiver**; code retirement is not gated. The controls were 080927Z FAIL and 082829Z FAIL (one flip each; the filing floor held; worldtree-dev calls it undecided). **Mine, 20260930T090608Z, PASSED (user median 0.83)**, so the count is 1 of 3 if the FAILs reset it. ⚠ It started by accident, from a test meant to be `--dry-run`, and matched the plan anyway.
|
||||||
|
- **⚠ U11b STEP 5 IS MINE, AUTO-TRIGGERED (Prime ruled 0320 via worldtree-dev, thread `01M3SGEQDRQD7DWBVT4K73FAHP`):** once the THIRD CONSECUTIVE PASS at off lands (infra-hermes copies me on every verdict; 20260930T090608Z was 1 of 3; a FAIL restarts the count), I re-run the H2 global agent-self=0 check (asked worldtree-dev for the exact procedure; do not improvise it), then delete on BOTH instances, using literal paths only: `agents/{forseti,lofn,mimir}/memory/<agent>.chroma` + `memory/context_promotion`. Default: stop the api, delete, start it (asked worldtree-dev whether b192 re-creates the dir). Then send worldtree-dev the stamp, because b193 ships after it. The 377 parked rows go with it. The archive stays under its rev 1.2 rules. After b193: remove the retired config keys at my pace.
|
||||||
- **U11b prep (worldtree-dev asks, read-only, answered 0215):** /embed usage from the Skuld ledger (`phase_name='embed'`, which carries no caller identity): demo 436 calls, 07-29..08-31, none since; personal 5 calls, 07-30..08-05. ⚠ Demo's whole Skuld ledger has been idle since 09-14. Legacy data: demo 3 chroma ≈2.1M plus context_promotion 224K; personal ≈26M plus 6.5M (1,413 JSONL + ledger.db). Both are in the `*_worldtree-state` volumes. **ARCHIVE DONE 2026-09-30 0957 PT (worldtree-dev GO), restore drill passed.** The copies:
|
- **U11b prep (worldtree-dev asks, read-only, answered 0215):** /embed usage from the Skuld ledger (`phase_name='embed'`, which carries no caller identity): demo 436 calls, 07-29..08-31, none since; personal 5 calls, 07-30..08-05. ⚠ Demo's whole Skuld ledger has been idle since 09-14. Legacy data: demo 3 chroma ≈2.1M plus context_promotion 224K; personal ≈26M plus 6.5M (1,413 JSONL + ledger.db). Both are in the `*_worldtree-state` volumes. **ARCHIVE DONE 2026-09-30 0957 PT (worldtree-dev GO), restore drill passed.** The copies:
|
||||||
- corviduo-dev `/var/lib/wt-legacy-archive/` (root 0700, UNENCRYPTED): demo 51 files and personal 1,426 files, each as tar.zst plus per-file sha256 plus MANIFEST.txt.
|
- corviduo-dev `/var/lib/wt-legacy-archive/` (root 0700, UNENCRYPTED): demo 51 files and personal 1,426 files, each as tar.zst plus per-file sha256 plus MANIFEST.txt.
|
||||||
- A DEDICATED restic repo, rest-server-nh3 `/nh3-dev/wt-legacy-archive/`, snapshot `98dc64e0`, with its password vaulted at `nh3-dev/wt-legacy-archive/restic-password`. Its URL is the vaulted `nh3-dev/etc/restic/repository` plus `wt-legacy-archive/`. It is mirrored to ana-nas at 05:00.
|
- A DEDICATED restic repo, rest-server-nh3 `/nh3-dev/wt-legacy-archive/`, snapshot `98dc64e0`, with its password vaulted at `nh3-dev/wt-legacy-archive/restic-password`. Its URL is the vaulted `nh3-dev/etc/restic/repository` plus `wt-legacy-archive/`. It is mirrored to ana-nas at 05:00.
|
||||||
@@ -178,6 +179,7 @@ _As of 2026-09-30 ~0120 PT._
|
|||||||
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
### SemIf on fv-ml1 GPU 1 (2026-09-27, Prime)
|
||||||
|
|
||||||
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
- **OFFLINE since 2026-09-30 0135 PT (Prime: "take semif offline for now; we'll optimize scriberr later").** Stopped with `docker compose stop`, not removed, to give scriberr back its GPU 1 room. Scriberr's Parakeet path hardcodes `--chunk-len 300`, and the attention memory grows with the square of the slice, so it needs over 6 GB; it hit CUDA OOM at 0124 on a 35-min file with ~6.7 GB free. Stopping SemIf moved GPU 1 from 91,052 to 81,806 MiB used. The same job re-run at 0137 finished clean: 35m17s of audio in 44 s. That is n=1, and the peak memory was not captured. **Deferred fix (Prime: later):** shorten scriberr's slice to ~120 s in our local build, then SemIf can come back. Embedding cards were ruled out: esh-ml1 has ~4.4 GB free and nh3-ml1 ~5.1 GB. A replacement bench (brokkr's Jev candidates) is running on GPU 3 under a separate harness.
|
||||||
|
- **⚠ Prime ~0840 2026-09-30: "replace semif with intern-decision now", on GPU 1, with Scriberr fixed alongside (his pick).** A background agent is building `intern-decision-serve` (semif-compatible API, :8033, `stacks/intern-decision`, token `intern-decision/api-token`) under a hard cap: total footprint ≤ 10,300 MiB. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: 16,081 free = 10,300 (intern-decision) + 5,496 (Scriberr) + 285 spare. The semif stack stays stopped as the rollback.
|
||||||
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. About 33 GB of candidate weights stay on fv-ml1 `/tank/aimodels/huggingface/hub` pending Prime. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
|
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. About 33 GB of candidate weights stay on fv-ml1 `/tank/aimodels/huggingface/hub` pending Prime. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
|
||||||
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
|
||||||
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +
|
||||||
|
|||||||
Reference in New Issue
Block a user