feat(augaman): v0.1.3 on esh-ml1 and fv-ml1; after-bench: GPU ~3x faster, CPU mode regressed
Both hosts are rebuilt from tag v0.1.3 (one ONNX session per detector canvas) and redeployed. pytest -m gpu tests/vision passes 3/3 on each card. Server-side for one face, same harness as the v0.1.2 baseline: - esh-ml1 GPU 144 -> 48 ms - fv-ml1 GPU 75 -> 27 ms End to end from nh3-dev: 101.6 and 73.8 ms. CPU mode got slower on every CPU target: fv-ml1 cpuset 0-5 went 152 -> 205 ms with a face, and the no-face frame roughly doubled. That is well outside the run-to-run spread. The suspected cause (not measured) is per-session ORT thread pools spinning. Reported to augaman-dev. Neither deployment uses CPU mode. On esh-ml1 the dependency layer missed the build cache and the rootfs touched 90% until the v0.1.2 image was removed. fv-ml1's build hit the cache, and the exported requirements are identical, so the stack README now says to check disk before building on esh-ml1.
This commit is contained in:
@@ -1,8 +1,42 @@
|
||||
# augaman speed bench: CPU vs esh-ml1 GPU vs fv-ml1 GPU
|
||||
|
||||
Asked by Prime on 2026-09-27. This is the **v0.1.2 "before"** baseline. v0.1.3 gives
|
||||
each detector canvas its own ONNX session, removing the ~90 ms CUDA shape-switching
|
||||
cost that comfy-dev found. The same harness is to be re-run on v0.1.3 for the "after".
|
||||
Asked by Prime on 2026-09-27. Two passes of the same harness:
|
||||
**v0.1.2 (before, 0021–0026 PT)** and **v0.1.3 (after, 0106–0112 PT)**. v0.1.3 gives
|
||||
each detector canvas its own ONNX session, which removes the ~90 ms CUDA
|
||||
shape-switching cost that comfy-dev found.
|
||||
|
||||
## Headline: server-side ms per frame, one face (a) / no face (c)
|
||||
|
||||
| target | v0.1.2 | **v0.1.3** | change |
|
||||
|---|---|---|---|
|
||||
| esh-ml1 GPU (RTX 2000E) | 143.8 / 136.2 | **48.0 / 37.2** | ~3.0× faster |
|
||||
| fv-ml1 GPU1 (RTX PRO 6000) | 75.0 / 69.5 | **27.3 / 22.9** | ~2.8× faster |
|
||||
| fv-ml1 CPU, cpuset 0-5 | 152.4 / 72.3 | **205.1 / 133.2** | ⚠ ~1.35× / 1.8× SLOWER |
|
||||
| fv-ml1 CPU, cpuset 0-23 | 155.9 / 61.9 | **204.6 / 124.0** | ⚠ SLOWER |
|
||||
| esh-ml1 CPU, 6 LXC threads | 887.6 / 159.4 | **1105.1 / 395.4** | ⚠ SLOWER |
|
||||
|
||||
End to end from nh3-dev, frame (a) p50: esh-ml1 GPU **101.6 ms** (was 196.5) and fv-ml1
|
||||
GPU **73.8 ms** (was 126.6). The /health floor is ~27 ms for both.
|
||||
|
||||
**The GPU gain is real, and CPU mode regressed.** The CPU slowdown sits far outside the
|
||||
run-to-run spread: fv CPU6 (c) ran 158–166 ms against 111–113 before. It shows on
|
||||
every CPU target and hits the no-face frame hardest, which points at the detector
|
||||
sessions. Hypothesis, NOT measured: each of the three ORT sessions now has its own
|
||||
intra-op thread pool, whose threads spin while another session runs, so the CPU
|
||||
oversubscribes. The candidate fixes are `session.intra_op.allow_spinning=0` or a shared
|
||||
global thread pool. Reported to augaman-dev. Neither deployment runs in CPU mode, so
|
||||
nothing live is affected.
|
||||
|
||||
VRAM after the bench (nvidia-smi): esh-ml1 736 MiB (514 on v0.1.2), fv-ml1 1828 MiB
|
||||
(1264 on v0.1.2). `pytest -m gpu tests/vision` passes 3/3 on both cards on v0.1.3
|
||||
(three sessions proven: detector@128, detector@640, embedder). Raw data:
|
||||
`rows-2026-09-27-v0.1.3.json`, and GPU utilisation samples every ~15 s in
|
||||
`gpu-util-2026-09-27-v0.1.3.log` (mostly 0%, with peaks of 21% on esh-ml1 and 31%
|
||||
on fv-ml1 GPU 1; fv-ml1 GPU 2 hit 99% once from another seat's traffic).
|
||||
|
||||
---
|
||||
|
||||
# v0.1.2 baseline detail
|
||||
|
||||
## Harness (it is part of the number)
|
||||
|
||||
|
||||
Reference in New Issue
Block a user