memory: snapshot — semif live + averaging spike (build 0.1.3 + fast-kernel trial next), restic creds out of units on all 8 hosts, infra-ops on vm-esh-nas, augaman fv-ml1 instance removed; 4 foot-guns

This commit is contained in:
vh
2026-09-27 02:55:39 -07:00
parent 739aa03123
commit d7ad235365
4 changed files with 135 additions and 10 deletions
@@ -0,0 +1,25 @@
# `[2026-09-27]` restic: repository URL out of the systemd units on all 8 hosts, and vaulted (Prime)
`resticprofile schedule` copies `env-file` values into the generated units under /etc/systemd/system
(0644). So `RESTIC_REPOSITORY`, including the rest-server basic-auth password, was readable by any local
user on every env-file host (found on esh-docker-vm; `systemctl cat` works without sudo). Blast radius:
the password allows reading the encrypted blobs and appending to one repo, and nothing more (rest-server
is `--append-only`, and the passphrase is a separate file).
**Fix** (`6e203dc`, `30f2c97`): `playbooks/restic-repository-file.yaml` does four things:
- derives `/etc/restic/repository` (root 0400) from `restic.env`;
- uploads the profile switched to `repository-file`, guarded by the live profile's pre-change sha;
- runs `cat config` through the new profile;
- re-schedules, then verifies the units exist and contain no `rest:http`.
Applied to ana-docker, fv-ml1 (`configs/restic/ana-ml2`), esh-docker-vm, esh-vm-db, irv-ml1, nh3-dev,
nh3-docker, and vm-esh-nas once Prime had bootstrapped infra-ops there. esh-ml1 was built that way. An
independent check across all 8 found 0 leaking units. nh3-docker's scheduled unit ran a real backup
(`a29b889d`). Every host's URL and passphrase are vaulted as `<host>/etc/restic/{repository,password}`.
- `restic.env` is **kept** (root 0600): the per-host READMEs and the freshness probe source it. **A rotation
must update the vault, `restic.env` and `repository`.**
- The rotation itself stays Prime's (backups runbook, Known gaps). The move stops the ongoing exposure but
does not un-expose what was readable.
- Before the edit, esh-docker-vm's and irv-ml1's live profiles had drifted ahead of the repo, and were
pulled in first (`c698751`).
@@ -0,0 +1,35 @@
# `[2026-09-27]` SemIf live on fv-ml1 GPU 1 (Prime)
**semif-serve 0.1.2 is live** at `http://10.251.50.54:8032` (`semif.fv.internal`), stack `stacks/semif`,
code and contract `services/semif-serve/` (`069725c`). It is a FastAPI wrapper around SemIf-OpenJev's
direct and shared torch scorers (MIT, pinned `23cf1f39fc9534fe81437200959b6dfc7106e45a`) on
Qwen3.5-4B `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (BF16, offline HF cache on `/tank`).
Hosting assessment: infra-hermes, `/tmp/semif.md` (2026-09-25).
Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4
GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt);
no consumer named yet, so every score is labelled uncalibrated.
**Two defects that only the card showed**, each fixed with a test:
- **0.1.1:** the OOM path raised `OutOfMemory(...) from exc` inside the `except` block. The chain kept the
torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed
allocated after the 503. The fix raises after the block, with no chain, and runs `gc.collect()`
before `empty_cache()`.
- **0.1.2:** a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1.
Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released
(`empty_cache`). It costs ~16 ms on a big request.
**Acceptance** (`services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`), against SemIf's
committed torch predictions on authored144:
- 142/144 same top choice; both misses are exact bf16 ties in our output;
- 144/144 identical prompt SHA-256;
- A-vs-A gap 0.0;
- negative control (rotated option descriptions) 14/144;
- shared vs direct 72/72;
- 21 binary criteria over one state in 159 ms.
The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.
**Build:** torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The
Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump
reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See [[2026-09-27-semif-order-averaging]].
@@ -0,0 +1,35 @@
# `[2026-09-27]` SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial
**Finding (Prime's probes):** SemIf's single-ordering answer leans toward whichever option is listed
first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty
call 0.760 when the order was reversed.
**Spike** (`739aa03`, `services/semif-serve/spike/`, no service change). SemIf authored144 +
perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one
shared request:
| method | accuracy |
|---|---|
| single ordering, as sent | 78.6% |
| expected single ordering | 80.0% |
| 3 rotations, log-mean | **87.7%** (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8) |
| all 6 orderings, log-mean | 88.1% |
Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%.
The service is deterministic, so the CI measures item sampling, not run noise. This is one task family
(evidence interpretation, 3 options), not our workload.
**Latency** (`spike/latency-2026-09-27.txt`, from nh3-dev, network floor 31 ms):
- `/decide` short: 71 ms end to end (38 ms server);
- 3 rotations shared: 113 ms (79 ms);
- 6 orderings: 137 ms (99 ms);
- a ~2k-token state: 210 ms (169 ms).
**Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.**
- Averaging design: opt-in `orderings: rotations|all` (all only ≤ 4 options); all orderings in one shared
batch; per-ordering SemIf results returned unchanged; a `combined` block (log-mean, top, agreement,
spread); orderings count toward `max_decisions`.
- Fast kernels: `causal_conv1d` + `flash-linear-attention` are missing, so transformers runs Qwen3.5's
reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.
Tracked in the in-flight SemIf section. See [[2026-09-27-semif-live-on-fv-ml1-gpu1]].