memory: true Jev is text-only (official docs); our real Jev gap is context (7,168 vs 64k tokens)

This commit is contained in:
vh
2026-09-30 13:09:44 -07:00
parent 1faadb45f6
commit e3dbb08d84
+1 -1
View File
@@ -190,7 +190,7 @@ _As of 2026-09-30 ~0120 PT._
- Acceptance on the live URL: 240/259 pooled and 79/84 Wyrd, with 0 of 560 rows changed against the bench.
- The semif container was REMOVED at 0949 via `compose down`; the image, files and token are kept (rollback in `stacks/semif/README.md`). There were no semif consumers to migrate.
- Still open: label ~50 real Wyrd/Cicada turns before trusting it in production (the card makes no contamination claim).
- **Jev API: the MODEL speaks it, the SERVICE does not.** Jev is TypeSafe's closed `jev-latest`: `POST /v1/systemone {state, model, questions:{id:{type noul|choice|score, instructions, criteria}}}` → `{answers:{id:{type, noul | choice+probabilities | probabilities}}, usage, model}`. The bench drove Intern-Decision natively through JevBench's `typesafe` adapter, 231 items × 4 repeats, all OK, so the engine's I/O matches that subset. intern-decision-serve exposes ONLY semif's `/decide` and `/decide/shared`, so a Jev client gets a 404. Adding `/v1/systemone` is a thin passthrough (the bench's 60-line wrapper is the seed); images would still be refused, because the vision tower is dropped for the GPU budget. **LIVE as 0.1.1 (infra-hermes ff552ab, deployed 1255 PT); infra-ops AUDIT PASSED at 1310.** My independent JevBench v1.2.16 typesafe run against the live endpoint: 202/231, hard 83/111, 0 row diffs against bench r1..r4 (the positive control, native vs drop-in, shows 11 diffs). Two low findings went back to infra-hermes: the contract's example shows `model` as an object while the wire uses a string, and `images: []` is wrongly refused. **Fixed in 0.1.2 (1866c00, live about 1303 PT), and my re-audit PASSED:** `images` [] and null are treated as absent, non-empty is 422, JevBench is still 202/231 with 0 diffs, and 124 tests pass. The rollback chain is 0.1.1 → 0.1.0. The pass line: JevBench v1.2.16's `typesafe` adapter against the live endpoint gives 202/231 (hard 83), with 0 row diffs against the bench ledgers. Also: >16 questions → 422 (no chunking), images → 422, peak ≤ 9,876 MiB, one inference thread only. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback.
- **Jev API: the MODEL speaks it, the SERVICE does not.** Jev is TypeSafe's closed `jev-latest`: `POST /v1/systemone {state, model, questions:{id:{type noul|choice|score, instructions, criteria}}}` → `{answers:{id:{type, noul | choice+probabilities | probabilities}}, usage, model}`. The bench drove Intern-Decision natively through JevBench's `typesafe` adapter, 231 items × 4 repeats, all OK, so the engine's I/O matches that subset. intern-decision-serve exposes ONLY semif's `/decide` and `/decide/shared`, so a Jev client gets a 404. Adding `/v1/systemone` is a thin passthrough (the bench's 60-line wrapper is the seed); images would still be refused, because the vision tower is dropped for the GPU budget. **LIVE as 0.1.1 (infra-hermes ff552ab, deployed 1255 PT); infra-ops AUDIT PASSED at 1310.** My independent JevBench v1.2.16 typesafe run against the live endpoint: 202/231, hard 83/111, 0 row diffs against bench r1..r4 (the positive control, native vs drop-in, shows 11 diffs). Two low findings went back to infra-hermes: the contract's example shows `model` as an object while the wire uses a string, and `images: []` is wrongly refused. **Fixed in 0.1.2 (1866c00, live about 1303 PT), and my re-audit PASSED:** `images` [] and null are treated as absent, non-empty is 422, JevBench is still 202/231 with 0 diffs, and 124 tests pass. The rollback chain is 0.1.1 → 0.1.0. **True Jev, per docs.typesafe.ai/models + /api (read 2026-09-30): TEXT ONLY** ("No image, audio, or video input"), so our images→422 matches it. Jev 1.13 allows 64k tokens per request (32k for state plus the longest question) and up to 255 options per choice; its errors are 422, 429 and 529. **Our real gap is context: MAX_TOKENS 7,168 against Jev's 64k,** set by the GPU 1 memory cap; a Jev client with a big state gets 422. The rest was probed live and matches: `usage` has input_tokens/output_tokens, instructions accepts an object or an array, and a 20-option choice returns 200. The pass line: JevBench v1.2.16's `typesafe` adapter against the live endpoint gives 202/231 (hard 83), with 0 row diffs against the bench ledgers. Also: >16 questions → 422 (no chunking), images → 422, peak ≤ 9,876 MiB, one inference thread only. **Scriberr fix LIVE 0900** (commit 0176ec0): `PARAKEET_CHUNK_THRESHOLD_SECS=120` + `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`, which drops the Parakeet peak from 9,384 to 5,496 MiB (n=3, deterministic). ⚠ My earlier claim that shorter slices cut memory ~6× was WRONG: a ~5.6 GB fixed floor dominates, and it is expandable_segments that cuts the fragmentation. GPU 1 budget: **15,442 MiB nvidia-smi Free** (my 16,081 was total minus used; the driver reserves ~640 MiB, which the build agent caught) = intern-decision at a 9.0 GiB cap (9,876 card peak; calls over ~7k tokens refused) + Scriberr 5,496 + 70 spare. The semif stack stays stopped as the rollback.
- **Jev replacement bench DONE 2026-09-30 0149–0456** (Prime's ask via brokkr, GPU 3, transient; the card is back to 2 MiB). **If SemIf is displaced, take Intern-Decision-4B on its own runtime.** It fits (9.7/10.3 GB) and is 1.5-2.3× faster (21 criteria in 88 vs 131 ms). It matches SemIf-with-rotations at ONE ordering (pooled +1.5, inside the ~4-pt floor) and is better on Wyrd. It is not a drop-in (new service + contract) and its card has no contamination statement. **JevBench rank does NOT transfer** to our sets: Plumb, the leader, is worse on Wyrd. The positive control reproduced exactly (SemIf 187/231, hard 0.613), and SemIf changed 0 labels across 4 restarts. The losing candidate weights (plumb-4b, JevK5 v0.2+v0.3, imajev-4b; about 24 GB) were DELETED on Prime's word at 1234 2026-09-30; Intern-Decision-4B is kept because it is live, and the pinned SHAs for a re-pull are in the bench doc. The doc is `docs/pfi/jev-candidates-bench-2026-09-30.md` (475d6d6); the deliverable went to brokkr on thread `01M3RPS5MW5CXMAHPFFS0DF39Y`.
- **Was LIVE: `semif-serve` 0.1.4** (was 0.1.3 until 1014 on 2026-09-27) at `http://10.251.50.54:8032` (`semif.fv.internal`), with order averaging
and the fast kernels. SemIf `23cf1f39`, Qwen3.5-4B `851bf6e8`, BF16; token `semif/api-token`. Code +