diff --git a/docs/pfi/augaman-speed-bench/README.md b/docs/pfi/augaman-speed-bench/README.md new file mode 100644 index 0000000..8c13b13 --- /dev/null +++ b/docs/pfi/augaman-speed-bench/README.md @@ -0,0 +1,79 @@ +# augaman speed bench: CPU vs esh-ml1 GPU vs fv-ml1 GPU + +Asked by Prime on 2026-09-27. This is the **v0.1.2 "before"** baseline. v0.1.3 gives +each detector canvas its own ONNX session, removing the ~90 ms CUDA shape-switching +cost that comfy-dev found. The same harness is to be re-run on v0.1.3 for the "after". + +## Harness (it is part of the number) + +- **Client:** `bench.py` on nh3-dev, sequential. Each request uses a **fresh + connection** (no keep-alive) and is timed around the whole POST. +- **Protocol:** 3 runs. Per (target, frame, run) there are 3 warm-up requests, + then 20 timed ones. Runs are interleaved across targets and frames, so drift + spreads over every condition. Time: 0021–0026 PT, 2026-09-27. +- **Server-side column:** the change in `augaman_pipeline_seconds{op="probe"}` + sum/count over the 20 timed requests. It is decode + detect + embed with no + network or HTTP, and it is **the device comparison**. It is a mean, not a p50. +- **Floor:** a `/health` round trip, same client. It came out 28–34 ms, higher than + augaman-dev's 16 ms because of the fresh connection per request. +- **Frames** (`make_frames.py`): (a) 1920x1080 grey canvas with one face; + (b) `obama_biden_2015.jpg` as-is, 960x1001, two faces; **(c) the same 1080p grey + canvas with no face**, a detector-only control. Every response was checked for + the expected face count, and **0 of 900 were wrong**. +- **Images:** `augaman:0.1.2`, built from the same tag on each host (separate + builds, same source). +- **Targets:** + +| target | device | notes | +|---|---|---| +| esh-ml1 GPU | RTX 2000E Ada 16 GB (50 W) | the production instance; shares the card with TEI + vllm-reward | +| fv-ml1 GPU1 | RTX PRO 6000 Blackwell Max-Q | the new second instance; shares GPU 1 with coder, erp and meromero seats | +| esh-ml1 CPU6 | the LXC's 6 host threads | throwaway `AUGAMAN_DEVICE=cpu` container. pve moved the LXC's cpuset between two reads, so it ran unpinned, with an affinity of 6 | +| fv-ml1 CPU6 | EPYC 9254, cpuset 0-5 | throwaway container; physical cores on socket 0 | +| fv-ml1 CPU24 | EPYC 9254, cpuset 0-23 | throwaway container; all physical cores of socket 0 | + +GPU utilisation read 0% on fv-ml1 GPU 1 and 0–9% on esh-ml1 across the samples +(0020–0021, 3 samples each), so the co-resident seats were idle. + +## Results (ms): median of the 3 run p50s, [min–max of the run p50s] + +| target | (a) 1 face, e2e | (b) 2 faces, e2e | (c) no face, e2e | **server (a)** | **server (b)** | **server (c)** | +|---|---|---|---|---|---|---| +| esh-ml1 GPU | 196.5 [195.0–196.6] | 195.1 [194.0–196.1] | 181.0 [177.6–182.9] | **143.8** | **138.7** | **136.2** | +| fv-ml1 GPU1 | 126.6 [121.8–128.3] | 122.7 [122.4–127.9] | 109.2 [108.7–111.0] | **75.0** | **73.9** | **69.5** | +| esh-ml1 CPU6 | 954.6 [792.8–1023.3] | 1325.6 [1306.9–1327.6] | 205.1 [202.1–228.0] | **887.6** | **1253.4** | **159.4** | +| fv-ml1 CPU6 | 203.6 [203.5–204.5] | 222.5 [216.0–236.4] | 111.6 [110.6–112.7] | **152.4** | **179.4** | **72.3** | +| fv-ml1 CPU24 | 208.1 [204.6–221.6] | 223.7 [219.8–237.1] | 101.1 [100.5–113.4] | **155.9** | **167.3** | **61.9** | + +The /health floor was 27.6–33.7 ms (median of run p50s) on every target. The raw +rows are in `rows-2026-09-27.json`, and `summarize.py` reproduces this table. + +## What it says + +- **fv-ml1's Blackwell is ~2× esh-ml1's Ada on v0.1.2:** 75 vs 144 ms server-side + for one face. End to end from nh3-dev it is 127 vs 197 ms. +- **On v0.1.2, a GPU buys little over a good CPU.** 6 EPYC cores do one face in + 152 ms, as fast as esh-ml1's GPU (144 ms). With no face, fv-ml1's CPU (62–72 ms) + matches fv-ml1's GPU (70 ms). The GPU pipeline is dominated by something other + than GPU compute, and the cards read 0–9% utilisation while it ran. That is + consistent with the detector's shape-switching cost that v0.1.3 removes. +- **esh-ml1's CPU is not a viable home:** 0.9 s for one face and 1.3 s for two. Its + embedder costs ~730 ms per face, against ~80 ms on 6 EPYC cores. That gap is much + larger than the detector gap (159 vs 72 ms), and its runs were the noisiest + (793–1023 ms), so treat the esh CPU number as indicative. It was also sharing its 6 + threads with TEI. +- **More EPYC threads do not help one frame:** CPU24 is roughly equal to CPU6 on faces. + +## Controls and sensitivity + +- **Repeats:** 3 runs × 20 per cell. The A-vs-A spread is the min–max above. +- **Noise floor:** the GPU cells spread by ≤6 ms across runs, and the fv CPU cells + by ≤20 ms. The esh-ml1 CPU (a) cell spread by 230 ms. Differences smaller than + those spreads are not findings. +- **Positive control (instrument agreement):** augaman-dev measured esh-ml1 (a) + independently at 186 ms with a 16 ms floor. This harness gets 196.5 ms with a + 28 ms floor. Minus each floor, 170 vs 169 ms: they agree. +- **Positive control (detects work):** frame (c) is fastest on every target, and + the gap from (a) to (c) is the embedding cost. It is visible on every device. +- **Negative control:** the face count was checked on all 900 responses, and none + were wrong. diff --git a/docs/pfi/augaman-speed-bench/bench.py b/docs/pfi/augaman-speed-bench/bench.py new file mode 100644 index 0000000..e4c6682 --- /dev/null +++ b/docs/pfi/augaman-speed-bench/bench.py @@ -0,0 +1,70 @@ +"""augaman /recognize latency bench, harness per augaman-dev (2026-09-27): +sequential from nh3-dev; per (target, frame, run): 3 warm-up requests, then 20 timed. +3 runs, interleaved across targets and frames so drift spreads over every condition. +Each request is a fresh connection (no keep-alive), timed around the whole POST. +Server-side share: delta of augaman_pipeline_seconds{op="probe"} sum/count over the +20 timed requests (a mean, not a p50). Every response is checked for the expected +face count; a mismatch invalidates the row. Floor: /health, 3 runs x 20, same client. + +Targets come from BENCH_TARGETS as JSON [{name, url, token_env}]; tokens are read +from the environment and never printed. +""" +import json, os, re, sys, time +import httpx + +FRAMES = [("a_1080p_1face", "frame_a_1080p_oneface.jpg", 1), + ("b_960x1001_2face", "obama_biden_2015.jpg", 2), + ("c_1080p_0face", "frame_c_1080p_noface.jpg", 0)] +RUNS, WARM, N = 3, 3, 20 +targets = json.loads(os.environ["BENCH_TARGETS"]) +data = {f: open(f, "rb").read() for _, f, _ in FRAMES} +PROBE = re.compile(r'^augaman_pipeline_seconds_(sum|count)\{[^}]*op="probe"[^}]*\}\s+([0-9.eE+-]+)$', re.M) + + +def probe_stats(url): + t = httpx.get(url + "/metrics", timeout=10).text + vals = {k: float(v) for k, v in PROBE.findall(t)} + return vals.get("sum", float("nan")), vals.get("count", float("nan")) + + +def recognize(url, tok, fname): + t0 = time.perf_counter() + r = httpx.post(url + "/recognize", headers={"Authorization": f"Bearer {tok}"}, + files={"frame": (fname, data[fname], "image/jpeg")}, timeout=120) + dt = (time.perf_counter() - t0) * 1000 + r.raise_for_status() + return dt, len(r.json()["faces"]) + + +def pct(xs, p): + xs = sorted(xs); k = (len(xs) - 1) * p; f = int(k); c = min(f + 1, len(xs) - 1) + return xs[f] + (xs[c] - xs[f]) * (k - f) + + +rows = [] +for run in range(1, RUNS + 1): + for t in targets: + url, tok = t["url"], os.environ[t["token_env"]] + # floor + h = [] + for _ in range(N): + t0 = time.perf_counter(); httpx.get(url + "/health", timeout=10).raise_for_status() + h.append((time.perf_counter() - t0) * 1000) + rows.append({"target": t["name"], "frame": "health_floor", "run": run, + "p50": pct(h, .5), "p90": pct(h, .9), "server_probe_mean": None, "bad": 0}) + for label, fname, want in FRAMES: + for _ in range(WARM): + recognize(url, tok, fname) + s0, c0 = probe_stats(url) + lat, bad = [], 0 + for _ in range(N): + dt, faces = recognize(url, tok, fname) + lat.append(dt); bad += faces != want + s1, c1 = probe_stats(url) + srv = (s1 - s0) / (c1 - c0) * 1000 if c1 > c0 else None + rows.append({"target": t["name"], "frame": label, "run": run, "p50": pct(lat, .5), + "p90": pct(lat, .9), "server_probe_mean": srv, "bad": bad}) + print(f"run{run} {t['name']:<16} {label:<17} p50 {pct(lat,.5):8.1f} p90 {pct(lat,.9):8.1f}" + f" srv {srv if srv is None else round(srv,1)} bad {bad}", flush=True) + +json.dump(rows, open(sys.argv[1] if len(sys.argv) > 1 else "bench_rows.json", "w"), indent=1) diff --git a/docs/pfi/augaman-speed-bench/make_frames.py b/docs/pfi/augaman-speed-bench/make_frames.py new file mode 100644 index 0000000..2da374c --- /dev/null +++ b/docs/pfi/augaman-speed-bench/make_frames.py @@ -0,0 +1,14 @@ +"""The bench frames, per augaman-dev's recipe (2026-09-27). Run in a directory holding +obama_2012.jpg and obama_biden_2015.jpg from pfi/augaman tests/fixtures/faces/ (public domain). + uv run --with pillow python make_frames.py +""" +from PIL import Image + +o = Image.open("obama_2012.jpg").convert("RGB") +h = 700 +o = o.resize((round(o.width * h / o.height), h), Image.LANCZOS) +a = Image.new("RGB", (1920, 1080), (90, 90, 90)) +a.paste(o, (700, 190)) +a.save("frame_a_1080p_oneface.jpg", quality=90) # (a) one face +Image.new("RGB", (1920, 1080), (90, 90, 90)).save("frame_c_1080p_noface.jpg", quality=90) # (c) control +# (b) is obama_biden_2015.jpg as-is: 960x1001, two faces diff --git a/docs/pfi/augaman-speed-bench/rows-2026-09-27.json b/docs/pfi/augaman-speed-bench/rows-2026-09-27.json new file mode 100644 index 0000000..1ea8416 --- /dev/null +++ b/docs/pfi/augaman-speed-bench/rows-2026-09-27.json @@ -0,0 +1,542 @@ +[ + { + "target": "esh-ml1 GPU", + "frame": "health_floor", + "run": 1, + "p50": 28.755177503626328, + "p90": 35.665129897824954, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "a_1080p_1face", + "run": 1, + "p50": 195.03028949839063, + "p90": 200.67508020147218, + "server_probe_mean": 143.80537764518522, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "b_960x1001_2face", + "run": 1, + "p50": 195.12132549425587, + "p90": 208.15129229886224, + "server_probe_mean": 139.3810058987583, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "c_1080p_0face", + "run": 1, + "p50": 182.8995855030371, + "p90": 191.65247399650985, + "server_probe_mean": 143.4166395556531, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "health_floor", + "run": 1, + "p50": 33.661384499282576, + "p90": 38.86325339844916, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "a_1080p_1face", + "run": 1, + "p50": 128.3326149932691, + "p90": 140.80376060592243, + "server_probe_mean": 74.98760224552825, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "b_960x1001_2face", + "run": 1, + "p50": 127.89153849735158, + "p90": 133.6793933995068, + "server_probe_mean": 74.19702970655635, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "c_1080p_0face", + "run": 1, + "p50": 111.00122199422913, + "p90": 117.57378029578831, + "server_probe_mean": 69.54348748549819, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "health_floor", + "run": 1, + "p50": 30.578960999264382, + "p90": 35.473932592140045, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 1, + "p50": 954.5748219970847, + "p90": 1020.2096714929212, + "server_probe_mean": 887.6064898984623, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 1, + "p50": 1327.6250180060742, + "p90": 1416.4716226019664, + "server_probe_mean": 1270.2313813992077, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 1, + "p50": 202.1412319954834, + "p90": 210.3825692931423, + "server_probe_mean": 159.44489179673837, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "health_floor", + "run": 1, + "p50": 27.60758400108898, + "p90": 29.759886700776406, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 1, + "p50": 203.55955500417622, + "p90": 219.8045792974881, + "server_probe_mean": 156.9333346793428, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 1, + "p50": 236.35765750077553, + "p90": 282.70697479892993, + "server_probe_mean": 188.06166611611843, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 1, + "p50": 110.56056749657728, + "p90": 114.69427619013004, + "server_probe_mean": 72.52865270711482, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "health_floor", + "run": 1, + "p50": 31.095413498405833, + "p90": 35.30468289682176, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "a_1080p_1face", + "run": 1, + "p50": 204.55308200325817, + "p90": 235.17365920706652, + "server_probe_mean": 159.80161383049563, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "b_960x1001_2face", + "run": 1, + "p50": 223.66256099485327, + "p90": 244.09297520178376, + "server_probe_mean": 171.44956310512498, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "c_1080p_0face", + "run": 1, + "p50": 101.08870249678148, + "p90": 114.38059099891689, + "server_probe_mean": 61.86515123117715, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "health_floor", + "run": 2, + "p50": 27.849696998600848, + "p90": 30.98080779891462, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "a_1080p_1face", + "run": 2, + "p50": 196.5537174983183, + "p90": 208.61720069660805, + "server_probe_mean": 144.78365819668397, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "b_960x1001_2face", + "run": 2, + "p50": 196.06467850098852, + "p90": 210.13995000248542, + "server_probe_mean": 138.73773735249415, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "c_1080p_0face", + "run": 2, + "p50": 180.95635149802547, + "p90": 188.9815433023614, + "server_probe_mean": 136.1989213983179, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "health_floor", + "run": 2, + "p50": 27.323893991706427, + "p90": 31.146193096356004, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "a_1080p_1face", + "run": 2, + "p50": 126.6464070067741, + "p90": 133.9685394996195, + "server_probe_mean": 74.96715487213805, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "b_960x1001_2face", + "run": 2, + "p50": 122.7374875015812, + "p90": 126.5989876977983, + "server_probe_mean": 73.87235894566402, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "c_1080p_0face", + "run": 2, + "p50": 108.72362100053579, + "p90": 114.4327986999997, + "server_probe_mean": 69.35211137169972, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "health_floor", + "run": 2, + "p50": 26.6998285005684, + "p90": 31.021830906684055, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 2, + "p50": 792.7710255025886, + "p90": 1075.1574820023964, + "server_probe_mean": 795.2748078998411, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 2, + "p50": 1325.6497279944597, + "p90": 1390.5041895020986, + "server_probe_mean": 1253.4432460510288, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 2, + "p50": 205.08419950056123, + "p90": 223.07179619820093, + "server_probe_mean": 168.24462699732976, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "health_floor", + "run": 2, + "p50": 27.37445900129387, + "p90": 32.55561870028032, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 2, + "p50": 204.45054300216725, + "p90": 220.60656060639303, + "server_probe_mean": 152.42901378078386, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 2, + "p50": 216.0444700057269, + "p90": 276.679380304995, + "server_probe_mean": 179.3103247997351, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 2, + "p50": 111.58617900218815, + "p90": 116.30833300296217, + "server_probe_mean": 72.20571864163503, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "health_floor", + "run": 2, + "p50": 27.235380999627523, + "p90": 29.255447497416764, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "a_1080p_1face", + "run": 2, + "p50": 208.08351850428153, + "p90": 221.3218400007463, + "server_probe_mean": 150.44436159078032, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "b_960x1001_2face", + "run": 2, + "p50": 219.75871399627067, + "p90": 259.07196740736254, + "server_probe_mean": 166.02391808992252, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "c_1080p_0face", + "run": 2, + "p50": 100.5097349989228, + "p90": 107.63760179688688, + "server_probe_mean": 59.30270317476243, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "health_floor", + "run": 3, + "p50": 26.72893099952489, + "p90": 31.63200489798328, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "a_1080p_1face", + "run": 3, + "p50": 196.49402699724305, + "p90": 204.47104229824618, + "server_probe_mean": 142.17575855000177, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "b_960x1001_2face", + "run": 3, + "p50": 194.01365749945398, + "p90": 203.67266629764345, + "server_probe_mean": 138.33359165437287, + "bad": 0 + }, + { + "target": "esh-ml1 GPU", + "frame": "c_1080p_0face", + "run": 3, + "p50": 177.55984649556922, + "p90": 184.63658349210164, + "server_probe_mean": 135.46126945293508, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "health_floor", + "run": 3, + "p50": 34.47381099977065, + "p90": 37.53316599468235, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "a_1080p_1face", + "run": 3, + "p50": 121.84531000093557, + "p90": 126.15792770229746, + "server_probe_mean": 75.25073911529034, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "b_960x1001_2face", + "run": 3, + "p50": 122.41429199639242, + "p90": 125.64870739151957, + "server_probe_mean": 73.74720136867836, + "bad": 0 + }, + { + "target": "fv-ml1 GPU1", + "frame": "c_1080p_0face", + "run": 3, + "p50": 109.15258450404508, + "p90": 114.19055700243916, + "server_probe_mean": 69.49142985977232, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "health_floor", + "run": 3, + "p50": 27.56448650325183, + "p90": 34.92588359804358, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 3, + "p50": 1023.3039910017396, + "p90": 1191.5021550987149, + "server_probe_mean": 939.7293467947748, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 3, + "p50": 1306.9450249968213, + "p90": 1427.1712208035751, + "server_probe_mean": 1251.0824149969267, + "bad": 0 + }, + { + "target": "esh-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 3, + "p50": 228.02160500577884, + "p90": 261.2633642929723, + "server_probe_mean": 153.0728507990716, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "health_floor", + "run": 3, + "p50": 53.94737450114917, + "p90": 107.14866449852707, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "a_1080p_1face", + "run": 3, + "p50": 203.5052149949479, + "p90": 222.22570729936706, + "server_probe_mean": 147.21799310063943, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "b_960x1001_2face", + "run": 3, + "p50": 222.4783574929461, + "p90": 276.977431504929, + "server_probe_mean": 179.41700990777463, + "bad": 0 + }, + { + "target": "fv-ml1 CPU6", + "frame": "c_1080p_0face", + "run": 3, + "p50": 112.70040849922225, + "p90": 125.06614489684581, + "server_probe_mean": 72.31316382531077, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "health_floor", + "run": 3, + "p50": 30.923282000003383, + "p90": 36.041059493436485, + "server_probe_mean": null, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "a_1080p_1face", + "run": 3, + "p50": 221.6059979982674, + "p90": 246.61138239607683, + "server_probe_mean": 155.86536910850555, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "b_960x1001_2face", + "run": 3, + "p50": 237.05747949861689, + "p90": 271.11399889254244, + "server_probe_mean": 167.3004158306867, + "bad": 0 + }, + { + "target": "fv-ml1 CPU24", + "frame": "c_1080p_0face", + "run": 3, + "p50": 113.37069750152295, + "p90": 157.2230929945363, + "server_probe_mean": 71.09609730541706, + "bad": 0 + } +] \ No newline at end of file diff --git a/docs/pfi/augaman-speed-bench/summarize.py b/docs/pfi/augaman-speed-bench/summarize.py new file mode 100644 index 0000000..d75653d --- /dev/null +++ b/docs/pfi/augaman-speed-bench/summarize.py @@ -0,0 +1,23 @@ +"""Summarise bench.py rows: per (target, frame), the median of the 3 run p50s with the +min-max of those run p50s (the A-vs-A spread, i.e. the noise floor), the median run p90, +and the median server-side probe mean. + python summarize.py rows-2026-09-27.json +""" +import json, statistics, sys +from collections import defaultdict + +rows = json.load(open(sys.argv[1])) +g = defaultdict(list) +for r in rows: + g[(r["target"], r["frame"])].append(r) +targets = list(dict.fromkeys(r["target"] for r in rows)) +frames = list(dict.fromkeys(r["frame"] for r in rows)) +print(f"{'target':<14} {'frame':<17} {'p50 med':>8} {'p50 run min-max':>16} {'p90 med':>8} {'server med':>10} bad") +for t in targets: + for f in frames: + rs = g[(t, f)] + p50s = [r["p50"] for r in rs] + srv = [r["server_probe_mean"] for r in rs if r["server_probe_mean"] is not None] + print(f"{t:<14} {f:<17} {statistics.median(p50s):8.1f} {min(p50s):7.1f}-{max(p50s):<8.1f}" + f" {statistics.median(r['p90'] for r in rs):8.1f}" + f" {(statistics.median(srv) if srv else float('nan')):10.1f} {sum(r['bad'] for r in rs)}") diff --git a/persistent-memory.md b/persistent-memory.md index 6ff746a..e110e30 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -155,6 +155,11 @@ _As of 2026-09-26 ~1620 PT._ survived the recreate and was then deleted, so the gallery is empty and ready for real enrollments. `/recognize` p50 186 ms (1080p, one face). The detector-latency follow-up is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops. +- **Second instance on fv-ml1 GPU 1 (Prime, 2026-09-27)**, `GPU_ID=1`, **fixtures-only, no + gallery backup**, separate gallery. Its role (bench-only vs future primary) is Prime's call. +- **Speed bench v0.1.2 "before"** (`docs/pfi/augaman-speed-bench/`), server-side one face: + esh GPU 144 ms, fv GPU 75 ms, fv CPU (6 cores) 152 ms, esh CPU 888 ms. **OWED: re-run on + v0.1.3** (detector gets one session per canvas, ~2× faster) when augaman-dev tags it. ### esh-ml1 diff --git a/servers/fv-ml1/README.md b/servers/fv-ml1/README.md index c3f65e5..f077de2 100644 --- a/servers/fv-ml1/README.md +++ b/servers/fv-ml1/README.md @@ -179,6 +179,16 @@ embed/rerank/reward trio. GPUs are pinned per container via **GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):** +> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`, +> `vllm-coder`, `vllm-erp-seat`, `vllm-meromero-rp` and now **`augaman`** (below). Read the +> host (`docker inspect … DeviceRequests`), not this table. +> +> **`augaman` :8040 (since 2026-09-27, Prime):** the second instance of the face-recognition service +> (`stacks/augaman`, `GPU_ID=1`), ~1.3 GB. **Fixtures-only: it has no gallery backup wired.** +> The primary instance, which holds the gallery and its backup, is on esh-ml1. Note that this host's restic copies +> `/var/lib/docker/volumes` raw, and that includes `augaman_gallery`: a live SQLite file, so the copy is +> not guaranteed consistent. That is acceptable for fixtures and not for real faces. + | Container | Port | Served model | Quant | Ctx | |-----------|------|--------------|-------|-----| | `vllm-granite` | 8004 | `granite-4.1-8b` — fleet summarizer/classifier | FP8 (compressed-tensors) | 131k | diff --git a/stacks/augaman/.env.example b/stacks/augaman/.env.example index 16a0cb8..ff9dc61 100644 --- a/stacks/augaman/.env.example +++ b/stacks/augaman/.env.example @@ -11,9 +11,19 @@ IMAGE=augaman:0.1.2 PORT=8040 HOST_IP=10.0.50.80 +# The host's card index for the GPU reservation (default 0). The container always +# sees its one card as index 0, so AUGAMAN_CUDA_DEVICE_ID never changes. +# esh-ml1: 0 (the only card) fv-ml1: 1 (the utility card, beside vllm-coder) +GPU_ID=0 +# Appended to the Homepage card name so the two instances are distinguishable. +# Note the leading space. esh-ml1 (the primary) leaves it empty. +# fv-ml1: CARD_SUFFIX=" (fv-ml1)" +CARD_SUFFIX= + # Where the backup CLI writes gallery.db. Owned 10001:10001 (the container user), -# mode 0700. It is the restic stage dir, so the backup run that copies the gallery here -# is the same run that ships it off-box (README "Backup"). +# mode 0700. On esh-ml1 it is the restic stage dir, so the backup run that copies the +# gallery here is the same run that ships it off-box (README "Backup"). +# fv-ml1 (fixtures-only, no scheduled gallery backup): /opt/docker/backup/augaman. BACKUP_DIR=/var/lib/restic/stage/augaman # >= 32 visible-ASCII characters. Source of truth is the vault: diff --git a/stacks/augaman/README.md b/stacks/augaman/README.md index 1b206f2..8ccdeb3 100644 --- a/stacks/augaman/README.md +++ b/stacks/augaman/README.md @@ -13,7 +13,14 @@ first, then re-mirror it here. | **Token** | `secret get augaman/api-token` (vault is the source of truth) | | **Image** | `augaman:`, built locally on esh-ml1 (below) | | **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) | -| **VRAM** | ~1.5 GB by design (embed batches capped at 16) | +| **VRAM** | ~0.5 GB on esh-ml1, ~1.3 GB on fv-ml1 (measured) | + +**Two instances (2026-09-27).** The **primary is on esh-ml1** (`GPU_ID=0`): it holds +the gallery and has the verified backup. A **second instance runs on fv-ml1** +(`GPU_ID=1`, the utility card beside `vllm-coder`; `CARD_SUFFIX=" (fv-ml1)"`; +`BACKUP_DIR=/opt/docker/backup/augaman`) at Prime's request. It has its own, +separate gallery, and it is **fixtures-only**: no gallery backup is wired there, +and the two galleries do not sync. Speed comparison: [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md). ## ⚠ Biometric data: backup gate diff --git a/stacks/augaman/compose.yaml b/stacks/augaman/compose.yaml index 23fcd58..66ddc54 100644 --- a/stacks/augaman/compose.yaml +++ b/stacks/augaman/compose.yaml @@ -1,7 +1,8 @@ -# augaman: the fleet's face-recognition service for Cicada (gitea pfi/augaman), on esh-ml1 -# (CT 110 on esh-pve, RTX 2000E Ada 16 GB). Enroll, recognize, verify; buffalo_l (SCRFD + -# ArcFace w600k_r50) on ONNX Runtime CUDA. Canonical copy: this file in pfi/augaman; the -# eshpfi stack mirrors it as stacks/augaman. +# augaman: the fleet's face-recognition service for Cicada (gitea pfi/augaman). The primary +# instance is on esh-ml1 (CT 110 on esh-pve, RTX 2000E Ada 16 GB), with the gallery and the +# backup. A second, fixtures-only instance runs on fv-ml1 (GPU_ID=1, no backup). Enroll, +# recognize, verify; buffalo_l (SCRFD + ArcFace w600k_r50) on ONNX Runtime CUDA. Canonical copy: +# this file in pfi/augaman; the eshpfi stack mirrors it as stacks/augaman. # # ⚠ BIOMETRIC DATA. The gallery volume holds face embeddings and crops of household members. # - The live SQLite stays on the local named volume. Never NFS (the service refuses it). @@ -21,7 +22,10 @@ # variables below carry no AUGAMAN_ prefix and are never passed through wholesale (no env_file). # # .env (tunables): IMAGE, PORT (8040), BACKUP_DIR, HOST_IP (10.0.50.80), AUGAMAN_API_TOKEN -# (>= 32 visible-ASCII characters; the source of truth is the vault). +# (>= 32 visible-ASCII characters; the source of truth is the vault), GPU_ID (the host's card +# index, default 0), CARD_SUFFIX (appended to the Homepage name, e.g. " (fv-ml1)"). +# AUGAMAN_CUDA_DEVICE_ID stays "0" on every host: the reservation shows the container only the +# card GPU_ID names, and it sees that card as index 0. name: augaman @@ -47,7 +51,7 @@ services: reservations: devices: - driver: nvidia - device_ids: ["0"] + device_ids: ["${GPU_ID:-0}"] capabilities: [gpu] healthcheck: # 200 only when ready and not degraded; a 503 (degraded) fails the check. @@ -58,7 +62,7 @@ services: start_period: 180s labels: - homepage.group=AI - Eval & Retrieval - - homepage.name=augaman — face recognition + - homepage.name=augaman — face recognition${CARD_SUFFIX:-} - homepage.icon=mdi-face-recognition - homepage.description=Enroll, recognize, verify (buffalo_l on CUDA) for Cicada - homepage.href=http://${HOST_IP:-10.0.50.80}:${PORT:-8040}/health