memory: snapshot — semif live + averaging spike (build 0.1.3 + fast-kernel trial next), restic creds out of units on all 8 hosts, infra-ops on vm-esh-nas, augaman fv-ml1 instance removed; 4 foot-guns
This commit is contained in:
@@ -0,0 +1,25 @@
|
||||
# `[2026-09-27]` restic: repository URL out of the systemd units on all 8 hosts, and vaulted (Prime)
|
||||
|
||||
`resticprofile schedule` copies `env-file` values into the generated units under /etc/systemd/system
|
||||
(0644). So `RESTIC_REPOSITORY`, including the rest-server basic-auth password, was readable by any local
|
||||
user on every env-file host (found on esh-docker-vm; `systemctl cat` works without sudo). Blast radius:
|
||||
the password allows reading the encrypted blobs and appending to one repo, and nothing more (rest-server
|
||||
is `--append-only`, and the passphrase is a separate file).
|
||||
|
||||
**Fix** (`6e203dc`, `30f2c97`): `playbooks/restic-repository-file.yaml` does four things:
|
||||
- derives `/etc/restic/repository` (root 0400) from `restic.env`;
|
||||
- uploads the profile switched to `repository-file`, guarded by the live profile's pre-change sha;
|
||||
- runs `cat config` through the new profile;
|
||||
- re-schedules, then verifies the units exist and contain no `rest:http`.
|
||||
|
||||
Applied to ana-docker, fv-ml1 (`configs/restic/ana-ml2`), esh-docker-vm, esh-vm-db, irv-ml1, nh3-dev,
|
||||
nh3-docker, and vm-esh-nas once Prime had bootstrapped infra-ops there. esh-ml1 was built that way. An
|
||||
independent check across all 8 found 0 leaking units. nh3-docker's scheduled unit ran a real backup
|
||||
(`a29b889d`). Every host's URL and passphrase are vaulted as `<host>/etc/restic/{repository,password}`.
|
||||
|
||||
- `restic.env` is **kept** (root 0600): the per-host READMEs and the freshness probe source it. **A rotation
|
||||
must update the vault, `restic.env` and `repository`.**
|
||||
- The rotation itself stays Prime's (backups runbook, Known gaps). The move stops the ongoing exposure but
|
||||
does not un-expose what was readable.
|
||||
- Before the edit, esh-docker-vm's and irv-ml1's live profiles had drifted ahead of the repo, and were
|
||||
pulled in first (`c698751`).
|
||||
@@ -0,0 +1,35 @@
|
||||
# `[2026-09-27]` SemIf live on fv-ml1 GPU 1 (Prime)
|
||||
|
||||
**semif-serve 0.1.2 is live** at `http://10.251.50.54:8032` (`semif.fv.internal`), stack `stacks/semif`,
|
||||
code and contract `services/semif-serve/` (`069725c`). It is a FastAPI wrapper around SemIf-OpenJev's
|
||||
direct and shared torch scorers (MIT, pinned `23cf1f39fc9534fe81437200959b6dfc7106e45a`) on
|
||||
Qwen3.5-4B `851bf6e806efd8d0a36b00ddf55e13ccb7b8cd0a` (BF16, offline HF cache on `/tank`).
|
||||
Hosting assessment: infra-hermes, `/tmp/semif.md` (2026-09-25).
|
||||
|
||||
Prime's calls (2026-09-27): fv-ml1 GPU 1 under a hard VRAM cap (not the reserved-empty GPU 3, not a Q4
|
||||
GGUF on a 16 GB box); infra-ops builds it with a light process (short contract → TDD → heid bug hunt);
|
||||
no consumer named yet, so every score is labelled uncalibrated.
|
||||
|
||||
**Two defects that only the card showed**, each fixed with a test:
|
||||
- **0.1.1:** the OOM path raised `OutOfMemory(...) from exc` inside the `except` block. The chain kept the
|
||||
torch exception's traceback, and with it the failed call's frames and tensors, so 11.9 GiB stayed
|
||||
allocated after the 503. The fix raises after the block, with no chain, and runs `gc.collect()`
|
||||
before `empty_cache()`.
|
||||
- **0.1.2:** a large shared request left torch's cache holding 12.6 GB, which left scriberr 3.5 GB on GPU 1.
|
||||
Now, after each call, reserved memory above the post-warm-up baseline + 512 MiB is released
|
||||
(`empty_cache`). It costs ~16 ms on a big request.
|
||||
|
||||
**Acceptance** (`services/semif-serve/acceptance/result-2026-09-27-v0.1.2.json`), against SemIf's
|
||||
committed torch predictions on authored144:
|
||||
- 142/144 same top choice; both misses are exact bf16 ties in our output;
|
||||
- 144/144 identical prompt SHA-256;
|
||||
- A-vs-A gap 0.0;
|
||||
- negative control (rotated option descriptions) 14/144;
|
||||
- shared vs direct 72/72;
|
||||
- 21 binary criteria over one state in 159 ms.
|
||||
|
||||
The shared-mode capacity under 12 GiB is 52 decisions at a ~140-token prefix and 13 at ~3,900.
|
||||
|
||||
**Build:** torch 2.10.0+cu128 from the pytorch index (SemIf's own stack; sm_120 present). The
|
||||
Dockerfile installs dependencies from a manifest with the project version blanked, so a version bump
|
||||
reuses the ~4 GB torch layer (41 s rebuild, layer CACHED). See [[2026-09-27-semif-order-averaging]].
|
||||
@@ -0,0 +1,35 @@
|
||||
# `[2026-09-27]` SemIf order-averaging: spiked, Prime ruled build + fast-kernel trial
|
||||
|
||||
**Finding (Prime's probes):** SemIf's single-ordering answer leans toward whichever option is listed
|
||||
first on ambiguous inputs. For "It's 2AM and I'm bored", the top pick flipped from casual 0.685 to booty
|
||||
call 0.760 when the order was reversed.
|
||||
|
||||
**Spike** (`739aa03`, `services/semif-serve/spike/`, no service change). SemIf authored144 +
|
||||
perturbations108, 252 rows in 72 groups, 3 options each, with all 6 orderings of every row sent in one
|
||||
shared request:
|
||||
|
||||
| method | accuracy |
|
||||
|---|---|
|
||||
| single ordering, as sent | 78.6% |
|
||||
| expected single ordering | 80.0% |
|
||||
| 3 rotations, log-mean | **87.7%** (+9.1 pts, group-bootstrap 95% CI +4.7..+13.8) |
|
||||
| all 6 orderings, log-mean | 88.1% |
|
||||
|
||||
Rows where the rotations agree unanimously: 161 rows, 94.4% accurate. Split rows: 91 rows, 75.8%.
|
||||
The service is deterministic, so the CI measures item sampling, not run noise. This is one task family
|
||||
(evidence interpretation, 3 options), not our workload.
|
||||
|
||||
**Latency** (`spike/latency-2026-09-27.txt`, from nh3-dev, network floor 31 ms):
|
||||
- `/decide` short: 71 ms end to end (38 ms server);
|
||||
- 3 rotations shared: 113 ms (79 ms);
|
||||
- 6 orderings: 137 ms (99 ms);
|
||||
- a ~2k-token state: 210 ms (169 ms).
|
||||
|
||||
**Prime ruled (2026-09-27 ~0250): build it in, and try the fast kernels too.**
|
||||
- Averaging design: opt-in `orderings: rotations|all` (all only ≤ 4 options); all orderings in one shared
|
||||
batch; per-ordering SemIf results returned unchanged; a `combined` block (log-mean, top, agreement,
|
||||
spread); orderings count toward `max_decisions`.
|
||||
- Fast kernels: `causal_conv1d` + `flash-linear-attention` are missing, so transformers runs Qwen3.5's
|
||||
reference PyTorch paths. Adopt them only if the authored144 parity re-check holds.
|
||||
|
||||
Tracked in the in-flight SemIf section. See [[2026-09-27-semif-live-on-fv-ml1-gpu1]].
|
||||
+40
-10
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-09-26 ~1620 PT (nh3-pve AMT live; nh3-ml1 load-shared + 3 brokkr foundry seats; esh-matter (Matter server) live for HA; augaman repo + agent live; coder stays on fv-ml1; VibeVoice → Q8_0.)_
|
||||
_Last updated: 2026-09-27 ~0255 PT (augaman v0.1.3 live on esh-ml1 with a verified off-site gallery backup; restic creds out of systemd units on all 8 hosts + vaulted; infra-ops on vm-esh-nas; SemIf live on fv-ml1 GPU1, order-averaging spike done, 0.1.3 build + fast-kernel trial next; corviduo/svoperatingsystem created.)_
|
||||
|
||||
> **Always check for `/tmp/infra-ops-handoff.md`** — if it exists and its
|
||||
> `Written:` stamp is under **8 hours** old, read it (it carries the in-flight
|
||||
@@ -115,7 +115,7 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-09-26 ~1620 PT._
|
||||
_As of 2026-09-27 ~0255 PT._
|
||||
|
||||
### nh3-pve + nh3-ml1: post-visit, all live (2026-09-25/26)
|
||||
|
||||
@@ -149,10 +149,28 @@ _As of 2026-09-26 ~1620 PT._
|
||||
Code + contract: `services/semif-serve/`; stack `stacks/semif`. **No consumer yet.**
|
||||
- Acceptance: 142/144 top choice vs upstream (both misses are exact bf16 ties), 144/144 prompt hashes,
|
||||
deterministic, negative control 14/144. It holds 12 GiB hard with release after a burst (~8.7 GB resting).
|
||||
- **Pending:** the heid bug-hunt panel (thread `01M3H3F4RR7XBP90KQ3A39H4SX`); triage and fold.
|
||||
- Latency from nh3-dev: `/decide` short 71 ms end to end (38 ms server); 3 rotations 113 ms; ~2k-token
|
||||
state 210 ms (`services/semif-serve/spike/latency-2026-09-27.txt`).
|
||||
- **NEXT, in this order (Prime 0250):**
|
||||
1. **Build order-averaging into semif-serve as 0.1.3.** Spike (`739aa03`): 3 rotations, log-mean,
|
||||
take accuracy 78.6% → 87.7% (group-bootstrap 95% CI +4.7..+13.8 pts); unanimous rows are 94.4%
|
||||
accurate, split rows 75.8%. Design: opt-in `orderings: rotations|all` (all only ≤ 4 options),
|
||||
every ordering in one shared batch, per-ordering SemIf results returned unchanged + a
|
||||
`combined` block (log-mean probabilities, top, agreement, per-option spread). Orderings count
|
||||
toward `max_decisions`. Contract amendment → TDD → deploy → acceptance (authored144 accuracy
|
||||
with/without; 2AM/2PM agreement as controls).
|
||||
2. **Trial the fast kernels** (`causal_conv1d`, `flash-linear-attention`). Today transformers runs
|
||||
Qwen3.5's reference PyTorch paths and logs that they are "much slower". Adopt only if the
|
||||
authored144 parity re-check holds.
|
||||
3. **heid bug-hunt panel** on 0.1.2 is pending (thread `01M3H3F4RR7XBP90KQ3A39H4SX`). Triage and fold
|
||||
its findings into 0.1.3.
|
||||
- NVFP4 is not worth it: the SemIf scorer runs on transformers, not vLLM; quant noise lands on the scores
|
||||
and would need its own calibration; BF16 already fits. SemIf's own MLX 4-bit run moved authored144
|
||||
0.813 → 0.789 (one run each, so indicative only).
|
||||
- Prime's probes (single scenarios, not benchmarks): at 2AM vs 2PM, booty-call probability averaged over
|
||||
all orderings goes 0.42 → 0.05; a baby dragon reads life-changing 0.65 / world-changing 0.25 with only
|
||||
2/3 rotations agreeing (one ordering said 0.925), the lottery ticket reads life-changing unanimously,
|
||||
and the paperclip control reads trivial 0.999.
|
||||
|
||||
### restic: credential leak fixed (2026-09-27, Prime)
|
||||
|
||||
@@ -163,12 +181,15 @@ _As of 2026-09-26 ~1620 PT._
|
||||
- **vm-esh-nas done too** (Prime bootstrapped infra-ops there at uid 850, NOPASSWD). **All 8 hosts are clean
|
||||
and vaulted.**
|
||||
- The passwords were readable until today, so the rotation (Prime's) remains the real fix.
|
||||
- **Check:** 2026-09-28 0100 PT is the first nightly run on `repository-file` for seven of the hosts
|
||||
(nh3-docker was proven end to end by a manual run, snapshot `a29b889d`). The 0800 freshness check
|
||||
should be all-green.
|
||||
|
||||
### augaman: face recognition for Cicada
|
||||
|
||||
- **v0.1.2 LIVE on esh-ml1:8040** (2026-09-27; v0.1.1 first deployed 2026-09-26 2347 PT),
|
||||
`stacks/augaman`, healthy on CUDA, in `nvidia-smi`. Built on-box from a `git archive`
|
||||
of the tag. `pytest -m gpu tests/vision` 3/3 PASS on v0.1.2.
|
||||
- **v0.1.3 LIVE on esh-ml1:8040** (v0.1.1 first deployed 2026-09-26 2347 PT), `stacks/augaman`,
|
||||
healthy on CUDA, in `nvidia-smi`. Built on-box from a `git archive` of the tag.
|
||||
`pytest -m gpu tests/vision` 3/3 PASS on v0.1.3.
|
||||
- **Gallery backup WIRED + RESTORE-VERIFIED 2026-09-27** (Prime OK'd off-site): restic
|
||||
daily 0100 PT → rest-server-ana `esh-ml1/`, fail-closed hook, secrets vaulted under
|
||||
`esh-ml1/etc/restic/`. Identity-level restore matched the live gallery (canary,
|
||||
@@ -179,7 +200,7 @@ _As of 2026-09-26 ~1620 PT._
|
||||
is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops.
|
||||
- **The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (~0130 PT)** after the bench. esh-ml1 is the only
|
||||
instance: in the house, backed up, 48 ms per face on v0.1.3.
|
||||
- **v0.1.3 live on BOTH hosts (2026-09-27 ~0105 PT).** Speed bench (`docs/pfi/augaman-speed-bench/`),
|
||||
- **v0.1.3 went live on both hosts (2026-09-27 ~0105 PT); the fv-ml1 instance was removed later.** Speed bench (`docs/pfi/augaman-speed-bench/`),
|
||||
server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode
|
||||
REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to
|
||||
augaman-dev; the deployment does not use CPU mode.
|
||||
@@ -195,13 +216,19 @@ _As of 2026-09-26 ~1620 PT._
|
||||
|
||||
### Live threads
|
||||
|
||||
- git: `main` is ahead of origin by 2 at `2fdbac6`, plus this snapshot. Pushing
|
||||
- git: `main` is ahead of origin by 3 at `739aa03`, plus this snapshot. Pushing
|
||||
is Prime's call.
|
||||
- ESH has a single outside route (esh-scale on esh-pve). Noted, untracked.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-09-26]` **augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy.** → `persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md`
|
||||
- `[2026-09-27]` **SemIf LIVE on fv-ml1 GPU 1 (semif-serve 0.1.2, Prime):** wrapper + contract + 39 tests, 142/144 upstream parity, two card-only memory defects fixed. → `persistent-memory.d/2026-09-27-semif-live-on-fv-ml1-gpu1.md`
|
||||
- `[2026-09-27]` **SemIf order-averaging spiked (+9.1 pts accuracy, agreement = strong confidence signal); Prime ruled: build it in as 0.1.3 and trial the fast kernels.** Tracked in the in-flight SemIf section. → `persistent-memory.d/2026-09-27-semif-order-averaging.md`
|
||||
- `[2026-09-27]` **restic repository URL moved out of world-readable systemd units on all 8 hosts, and vaulted (Prime).** Rotation stays Prime's. → `persistent-memory.d/2026-09-27-restic-repository-file-fleetwide.md`
|
||||
- `[2026-09-27]` **infra-ops bootstrapped on vm-esh-nas by Prime** (the first account at the fleet-pinned uid/gid 850; `bootstrap-infra-ops-user.yaml` now pins 850 when free, `d775a01`).
|
||||
- `[2026-09-27]` **augaman: second instance on fv-ml1 benched, then REMOVED (Prime).** esh-ml1 on v0.1.3 does 48 ms/face in the house with the verified backup (`ad484c3`, bench `docs/pfi/augaman-speed-bench/`).
|
||||
- `[2026-09-27]` Created empty private repo `corviduo/svoperatingsystem` (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. `svos-dev` → the OS session; The High Seat's session is `highseat-dev` again.
|
||||
- `[2026-09-26]` **augaman: infra-ops owes the gallery backup + first deploy — DEFERRED until augaman-dev reaches deploy.** → `persistent-memory.d/2026-09-26-augaman-infra-ops-owes-the-gallery-backup-first-deploy.md` — `[2026-09-27]` **DONE:** deployed (v0.1.3 on esh-ml1), backup wired, identity-level restore verified (`d8f59a1`).
|
||||
- `[2026-09-26]` **`zellij-fleet@.service` installed, NOT enabled** (for svos-dev's seat_up; owns the fleet zellij server in its own cgroup). Enabling `@Claude` at boot is Prime's call when seat_up ships. Tracked: `0ad7799`, `services/zellij-fleet/README.md`.
|
||||
- `[2026-09-26]` **btusb blacklist on nh3-pve + esh-pve — proposed, awaiting Prime (low priority).** Every 6.8.12-4x boot oopses in btmtk (benign). Also noted for the next NH3 visit: nh3-pve runs ~20 °C warmer than esh-pve (workload-confounded; check airflow). Tracked: `servers/nh3-pve/README.md`.
|
||||
- `[2026-09-26]` **GPU-LXC Temperature alerts watch the GPU, not the host CPU** (`SENSORS=-coretemp_*,acpitz`); hypervisor CPU alerts at >95 °C on nh3-pve/esh-pve. The first nh3-ml1 alert was nh3-pve's CPU during vzdump. `6d901c0`.
|
||||
@@ -211,7 +238,6 @@ _As of 2026-09-26 ~1620 PT._
|
||||
- `[2026-09-25]` **pfi-gx10 AC-restore VALIDATED by Prime's AC pull; Homepage stays on esh-docker-vm (Prime) after the mmap_lock wedge reboot.** `servers/pfi-gx10/README.md`, `servers/esh-docker-vm/README.md`.
|
||||
- `[2026-09-26]` **VibeVoice ASR → Q8_0 (Prime).** WER on the 4 bundled LibriSpeech clips 3/69 → 2/69 (only the I'm/I am artifact left), ×3 identical; RTF 0.09–0.17; +1.1 GB VRAM (nh3-ml1 ~11.3/16 GB). Q4_K file removed.
|
||||
- `[2026-09-26]` **esh-matter LIVE: a Matter server (matter.js 1.4.0) on CT 111 @ 10.0.90.20, VLAN 90 only** — , for ha-dev (operator-approved, relayed). → `persistent-memory.d/2026-09-26-esh-matter-live-a-matter-server-matter-js-1-4-0-on-ct-111.md`
|
||||
- `[2026-09-27]` Created empty private repo `corviduo/svoperatingsystem` (SVOS, the operating system) as claude-bot, at brokkr-smithy-dev's relay of Prime's ruling. `svos-dev` → the OS session; The High Seat's session is `highseat-dev` again.
|
||||
- `[2026-09-26]` **Embed/rerank LOAD-SHARED across esh-ml1 + nh3-ml1 (Prime).** — Second deployments were added for qwen3-embedding and reranker (config) and for reranker-a3-bge-v2-m3 (DB… → `persistent-memory.d/2026-09-26-embed-rerank-load-shared-across-esh-ml1-nh3-ml1-prime.md`
|
||||
- `[2026-09-26]` **Two dataset-foundry utility seats LIVE on nh3-ml1 (brokkr; operator approval relayed):** — LFM2.5-VL-3B on llama.cpp `:8030` (gateway `lfm25-vl-3b`, LiteLLM restarted 36 s at 0039) and… → `persistent-memory.d/2026-09-26-two-dataset-foundry-utility-seats-live-on-nh3-ml1-brokkr.md`
|
||||
- `[2026-09-26]` **Coder seat STAYS on fv-ml1 (Prime).** — The nh3-ml1 copy gave the same quality (teacher-forced true-code logprob diff +0.008 ± 0.019) but ran ~5×… → `persistent-memory.d/2026-09-26-coder-seat-stays-on-fv-ml1-prime.md`
|
||||
@@ -414,6 +440,10 @@ _106 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-09-27]` **`uv export` → requirements.txt → `uv pip install` for a torch-from-cu128 project** — the export drops uv's per-package index routing; with `--extra-index-url … --index-strategy unsafe-best-match`, triton came from the pytorch index and failed the lock's hash. Keep `uv sync` and blank the project version in a deps-only stage instead (`services/semif-serve/Dockerfile`).
|
||||
- `[2026-09-27]` **Translating a CUDA OOM by raising inside `except` (or `from exc`)** — the chained exception's traceback pins the failed call's frames, and their GPU tensors (11.9 GiB after the 503). Raise after the block, unchained, `gc.collect()` before `empty_cache()`.
|
||||
- `[2026-09-27]` **`deploy-stack.sh --yes` / `dns-sync.py` fed a blind `y`** — the harness refuses a blind apply. Review with `echo n |` first, then apply with `echo y |`.
|
||||
- `[2026-09-27]` **Relying on the build cache surviving on esh-ml1** — the augaman v0.1.3 build missed its dependency layer there (fv-ml1 hit it; reason unexplained, suspect the `docker builder prune` runs) and the rootfs peaked at 90%. Check for ≥15 GB free, or remove the old image, before building there.
|
||||
- `[2026-09-26]` **Greedy exact-match as parity for a small generative seat (coder)** — same-host repeats matched only 30–38% (run B hit the prefix cache, so a different numeric path; near-tie… → `persistent-memory.d/2026-09-26-greedy-exact-match-as-parity-for-a-small-generative-seat.md`
|
||||
- `[2026-09-26]` **zsh traps in ad-hoc test loops:** `set -- $a` does NOT word-split, so alert POSTs went out with empty fields; and `local path=…` inside a function CLOBBERS `$PATH` (zsh's tied array), so the failover test ran nothing while TEI sat stopped. Use other names and `${=var}`.
|
||||
- `[2026-09-25]` **Concluding "not on the UDM" from a port table read 15 s after link-up** — UniFi polls (~60 s); AMT was on port 6 as Prime said. Wait a poll, and prefer a which-port-lists-the-MAC check. auto-memory `feedback_polled_stats_lag_the_event`.
|
||||
|
||||
Reference in New Issue
Block a user