memory: snapshot — semif live + averaging spike (build 0.1.3 + fast-kernel trial next), restic creds out of units on all 8 hosts, infra-ops on vm-esh-nas, augaman fv-ml1 instance removed; 4 foot-guns
This commit is contained in:
@@ -0,0 +1,25 @@
|
||||
# `[2026-09-27]` restic: repository URL out of the systemd units on all 8 hosts, and vaulted (Prime)
|
||||
|
||||
`resticprofile schedule` copies `env-file` values into the generated units under /etc/systemd/system
|
||||
(0644). So `RESTIC_REPOSITORY`, including the rest-server basic-auth password, was readable by any local
|
||||
user on every env-file host (found on esh-docker-vm; `systemctl cat` works without sudo). Blast radius:
|
||||
the password allows reading the encrypted blobs and appending to one repo, and nothing more (rest-server
|
||||
is `--append-only`, and the passphrase is a separate file).
|
||||
|
||||
**Fix** (`6e203dc`, `30f2c97`): `playbooks/restic-repository-file.yaml` does four things:
|
||||
- derives `/etc/restic/repository` (root 0400) from `restic.env`;
|
||||
- uploads the profile switched to `repository-file`, guarded by the live profile's pre-change sha;
|
||||
- runs `cat config` through the new profile;
|
||||
- re-schedules, then verifies the units exist and contain no `rest:http`.
|
||||
|
||||
Applied to ana-docker, fv-ml1 (`configs/restic/ana-ml2`), esh-docker-vm, esh-vm-db, irv-ml1, nh3-dev,
|
||||
nh3-docker, and vm-esh-nas once Prime had bootstrapped infra-ops there. esh-ml1 was built that way. An
|
||||
independent check across all 8 found 0 leaking units. nh3-docker's scheduled unit ran a real backup
|
||||
(`a29b889d`). Every host's URL and passphrase are vaulted as `<host>/etc/restic/{repository,password}`.
|
||||
|
||||
- `restic.env` is **kept** (root 0600): the per-host READMEs and the freshness probe source it. **A rotation
|
||||
must update the vault, `restic.env` and `repository`.**
|
||||
- The rotation itself stays Prime's (backups runbook, Known gaps). The move stops the ongoing exposure but
|
||||
does not un-expose what was readable.
|
||||
- Before the edit, esh-docker-vm's and irv-ml1's live profiles had drifted ahead of the repo, and were
|
||||
pulled in first (`c698751`).
|
||||
@@ -0,0 +1,35 @@
|
||||
# `[2026-09-27]` SemIf live on fv-ml1 GPU 1 (Prime)
|
||||
|
||||
**semif-serve 0.1.2 is live** at `http://10.251.50.54:8032` (`semif.fv.internal`), stack `stacks/semif`,
|
||||
code and contract `services/semif-serve/` (`069725c`). It is a FastAPI wrapper around SemIf-OpenJev's
|
||||
direct and shared torch scorers (MIT, pinned `23cf1f39fc9534fe81437200959b6dfc7106e45a`) on
|
||||
Qwen3.5-4B `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (BF16, offline HF cache on `/tank`).
|
||||
Hosting assessment: infra-hermes, `/tmp/semif.md` (2026-09-25).
|
||||
|
||||
Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4
|
||||
GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt);
|
||||
no consumer named yet, so every score is labelled uncalibrated.
|
||||
|
||||
**Two defects that only the card showed**, each fixed with a test:
|
||||
- **0.1.1:** the OOM path raised `OutOfMemory(...) from exc` inside the `except` block. The chain kept the
|
||||
torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed
|
||||
allocated after the 503. The fix raises after the block, with no chain, and runs `gc.collect()`
|
||||
before `empty_cache()`.
|
||||
- **0.1.2:** a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1.
|
||||
Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released
|
||||
(`empty_cache`). It costs ~16 ms on a big request.
|
||||
|
||||
**Acceptance** (`services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`), against SemIf's
|
||||
committed torch predictions on authored144:
|
||||
- 142/144 same top choice; both misses are exact bf16 ties in our output;
|
||||
- 144/144 identical prompt SHA-256;
|
||||
- A-vs-A gap 0.0;
|
||||
- negative control (rotated option descriptions) 14/144;
|
||||
- shared vs direct 72/72;
|
||||
- 21 binary criteria over one state in 159 ms.
|
||||
|
||||
The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.
|
||||
|
||||
**Build:** torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The
|
||||
Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump
|
||||
reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See [[2026-09-27-semif-order-averaging]].
|
||||
@@ -0,0 +1,35 @@
|
||||
# `[2026-09-27]` SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial
|
||||
|
||||
**Finding (Prime's probes):** SemIf's single-ordering answer leans toward whichever option is listed
|
||||
first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty
|
||||
call 0.760 when the order was reversed.
|
||||
|
||||
**Spike** (`739aa03`, `services/semif-serve/spike/`, no service change). SemIf authored144 +
|
||||
perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one
|
||||
shared request:
|
||||
|
||||
| method | accuracy |
|
||||
|---|---|
|
||||
| single ordering, as sent | 78.6% |
|
||||
| expected single ordering | 80.0% |
|
||||
| 3 rotations, log-mean | **87.7%** (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8) |
|
||||
| all 6 orderings, log-mean | 88.1% |
|
||||
|
||||
Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%.
|
||||
The service is deterministic, so the CI measures item sampling, not run noise. This is one task family
|
||||
(evidence interpretation, 3 options), not our workload.
|
||||
|
||||
**Latency** (`spike/latency-2026-09-27.txt`, from nh3-dev, network floor 31 ms):
|
||||
- `/decide` short: 71 ms end to end (38 ms server);
|
||||
- 3 rotations shared: 113 ms (79 ms);
|
||||
- 6 orderings: 137 ms (99 ms);
|
||||
- a ~2k-token state: 210 ms (169 ms).
|
||||
|
||||
**Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.**
|
||||
- Averaging design: opt-in `orderings: rotations|all` (all only ≤ 4 options); all orderings in one shared
|
||||
batch; per-ordering SemIf results returned unchanged; a `combined` block (log-mean, top, agreement,
|
||||
spread); orderings count toward `max_decisions`.
|
||||
- Fast kernels: `causal_conv1d` + `flash-linear-attention` are missing, so transformers runs Qwen3.5's
|
||||
reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.
|
||||
|
||||
Tracked in the in-flight SemIf section. See [[2026-09-27-semif-live-on-fv-ml1-gpu1]].
|
||||
Reference in New Issue
Block a user