Both hosts are rebuilt from tag v0.1.3 (one ONNX session per detector canvas) and redeployed. pytest -m gpu tests/vision passes 3/3 on each card. Server-side for one face, same harness as the v0.1.2 baseline: - esh-ml1 GPU 144 -> 48 ms - fv-ml1 GPU 75 -> 27 ms End to end from nh3-dev: 101.6 and 73.8 ms. CPU mode got slower on every CPU target: fv-ml1 cpuset 0-5 went 152 -> 205 ms with a face, and the no-face frame roughly doubled. That is well outside the run-to-run spread. The suspected cause (not measured) is per-session ORT thread pools spinning. Reported to augaman-dev. Neither deployment uses CPU mode. On esh-ml1 the dependency layer missed the build cache and the rootfs touched 90% until the v0.1.2 image was removed. fv-ml1's build hit the cache, and the exported requirements are identical, so the stack README now says to check disk before building on esh-ml1.
87 lines
4.7 KiB
Markdown
87 lines
4.7 KiB
Markdown
# augaman
|
||
|
||
**The fleet's face-recognition service for Cicada** (enroll, recognize, verify),
|
||
on **esh-ml1** (CT 110 on esh-pve, RTX 2000E Ada). buffalo_l (SCRFD detector +
|
||
ArcFace w600k_r50) on ONNX Runtime CUDA. Code and the canonical compose live in
|
||
gitea `pfi/augaman` (owner: `augaman-dev`); `compose.yaml` here is a verbatim
|
||
mirror of that repo's `deploy/compose.yaml` at the deployed tag. Change it there
|
||
first, then re-mirror it here.
|
||
|
||
| | |
|
||
|---|---|
|
||
| **URL** | `http://10.0.50.80:8040` (`/health` is unauthenticated; everything else needs the bearer token) |
|
||
| **Token** | `secret get augaman/api-token` (vault is the source of truth) |
|
||
| **Image** | `augaman:<version>`, built locally on esh-ml1 (below) |
|
||
| **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) |
|
||
| **VRAM** | ~0.5 GB on esh-ml1, ~1.3 GB on fv-ml1 (measured) |
|
||
|
||
**Two instances (2026-09-27).** The **primary is on esh-ml1** (`GPU_ID=0`): it holds
|
||
the gallery and has the verified backup. A **second instance runs on fv-ml1**
|
||
(`GPU_ID=1`, the utility card beside `vllm-coder`; `CARD_SUFFIX=" (fv-ml1)"`;
|
||
`BACKUP_DIR=/opt/docker/backup/augaman`) at Prime's request. It has its own,
|
||
separate gallery, and it is **fixtures-only**: no gallery backup is wired there,
|
||
and the two galleries do not sync. Speed comparison: [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md).
|
||
|
||
## ⚠ Biometric data: backup gate
|
||
|
||
The gallery holds face embeddings and crops of household members. esh-ml1 is
|
||
**outside vzdump**, so the gallery reaches backup only through the app's backup
|
||
CLI, which writes `gallery.db` (mode 0600) into `BACKUP_DIR` =
|
||
`/var/lib/restic/stage/augaman` (owned 10001:10001, mode 0700).
|
||
|
||
**Operator ruling: until a scheduled backup ships that file off-box AND one
|
||
restore has been verified (the restored copy reports the same identities), only
|
||
public-domain test fixtures may be enrolled. No household faces.**
|
||
|
||
Backup status (2026-09-27): **WIRED AND RESTORE-VERIFIED, so the gate is met.**
|
||
restic runs daily at 0100 PT to rest-server-ana, with a fail-closed pre-backup
|
||
hook ([`configs/restic/esh-ml1/`](../../configs/restic/esh-ml1/README.md)). The
|
||
restore was verified against augaman-dev's public-domain canary identity (1
|
||
identity, 3 samples): the copy restored from snapshot `fd3061a1` matched the
|
||
live gallery (identities, samples and embeddings, by digest), and
|
||
`integrity_check` returned ok.
|
||
|
||
```bash
|
||
docker exec augaman python -m augaman.gallery.backup --db /data/gallery.db --dest /backup/gallery.db
|
||
# one JSON line + exit 0, or "backup failed: ..." + exit 1
|
||
```
|
||
|
||
## Building
|
||
|
||
The image is built on esh-ml1 from the release tag's content. esh-ml1 has no
|
||
gitea credentials, so the source is shipped as a `git archive`:
|
||
|
||
```bash
|
||
# from nh3-dev, in a pfi/augaman checkout
|
||
git archive --format=tar vX.Y.Z | ssh esh-ml1 'sudo mkdir -p /opt/docker/src/augaman-vX.Y.Z && sudo tar -x -C /opt/docker/src/augaman-vX.Y.Z && sudo chown -R infra-ops:infra-ops /opt/docker/src'
|
||
# on esh-ml1
|
||
cd /opt/docker/src/augaman-vX.Y.Z && docker build -t augaman:X.Y.Z .
|
||
```
|
||
|
||
The build fetches the two pinned models and SHA-256-checks them; a mismatch fails
|
||
the build. v0.1.0 built in under 2 minutes cold; v0.1.1 was the first version
|
||
deployed (2026-09-26), then v0.1.2 and v0.1.3 (2026-09-27; v0.1.3 gives each
|
||
detector canvas its own session, about +190 MiB VRAM on esh-ml1).
|
||
|
||
⚠ **Disk:** a build that has to install the third-party packages takes ~8–11 GB
|
||
transiently (image ~5 GB plus the ~3 GB uv download cache). Beszel alerts at 85%.
|
||
Before v0.1.2, every version bump re-installed them, and the v0.1.1 build hit 90%.
|
||
From v0.1.2 the packages install from a layer keyed on `uv export
|
||
--no-emit-project`, so a bump that changes no dependency reuses that layer. Keep
|
||
the layer cache (`docker builder prune --filter type=exec.cachemount` drops only
|
||
the download cache); a full `docker builder prune` forces the next build to
|
||
re-install everything. ⚠ **The cache did not survive on esh-ml1 for v0.1.3** (the
|
||
deps layer re-ran for 1m43s, and the rootfs touched 90% until the old image was
|
||
removed). On fv-ml1 the same build was CACHED (49 s), and the exported requirements
|
||
were byte-identical across the two versions, so the layering itself works. The
|
||
esh-ml1 miss is unexplained; suspect the cache pruning done there after v0.1.2.
|
||
**On esh-ml1, check `df` before a build, and have ≥15 GB free or remove the old
|
||
image first.**
|
||
|
||
## Startup is fail-closed on CUDA
|
||
|
||
The service refuses to serve unless CUDA really runs every convolution (profiled
|
||
warmup, ORT CPU fallback disabled). `/health` reports backend `cuda` only then.
|
||
Allow up to 180 s (`start_period`). The real GPU confirmation is the augaman
|
||
process showing up in `nvidia-smi` on the host.
|