Files
esh-pfi-infrastructure/stacks/augaman/README.md
T
vh 317868dc7e feat(augaman): second, fixtures-only instance on fv-ml1 GPU 1; CPU vs GPU speed bench (v0.1.2 baseline)
Prime asked for augaman on fv-ml1's utility card, beside vllm-coder. Mirror
augaman-dev's f77164f compose, which parameterises the GPU reservation (GPU_ID,
default 0) and the Homepage card name (CARD_SUFFIX). esh-ml1's resolved config is
unchanged: same config hash, no recreate.

On fv-ml1: augaman:0.1.2 built on-box from the tag, GPU_ID=1, healthy on CUDA
at 1264 MiB, and pytest -m gpu tests/vision passes 3/3 on the Blackwell. It has
its own gallery and no gallery backup, so it is fixtures-only. The host's raw
restic copy of /var/lib/docker/volumes is not a consistent SQLite backup.

docs/pfi/augaman-speed-bench/ holds the harness (augaman-dev's recipe plus a
no-face control frame and a face-count check on every response), the raw rows
and the summary. Server-side, one face:
- esh-ml1 GPU 144 ms
- fv-ml1 GPU 75 ms
- fv-ml1 CPU on 6 cores 152 ms
- esh-ml1 CPU 888 ms
It agrees with augaman-dev's independent esh-ml1 measurement once each
harness's floor is subtracted. This is the before for v0.1.3's detector fix.
2026-09-27 00:28:20 -07:00

80 lines
4.1 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# augaman
**The fleet's face-recognition service for Cicada** (enroll, recognize, verify),
on **esh-ml1** (CT 110 on esh-pve, RTX 2000E Ada). buffalo_l (SCRFD detector +
ArcFace w600k_r50) on ONNX Runtime CUDA. Code and the canonical compose live in
gitea `pfi/augaman` (owner: `augaman-dev`); `compose.yaml` here is a verbatim
mirror of that repo's `deploy/compose.yaml` at the deployed tag. Change it there
first, then re-mirror it here.
| | |
|---|---|
| **URL** | `http://10.0.50.80:8040` (`/health` is unauthenticated; everything else needs the bearer token) |
| **Token** | `secret get augaman/api-token` (vault is the source of truth) |
| **Image** | `augaman:<version>`, built locally on esh-ml1 (below) |
| **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) |
| **VRAM** | ~0.5 GB on esh-ml1, ~1.3 GB on fv-ml1 (measured) |
**Two instances (2026-09-27).** The **primary is on esh-ml1** (`GPU_ID=0`): it holds
the gallery and has the verified backup. A **second instance runs on fv-ml1**
(`GPU_ID=1`, the utility card beside `vllm-coder`; `CARD_SUFFIX=" (fv-ml1)"`;
`BACKUP_DIR=/opt/docker/backup/augaman`) at Prime's request. It has its own,
separate gallery, and it is **fixtures-only**: no gallery backup is wired there,
and the two galleries do not sync. Speed comparison: [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md).
## ⚠ Biometric data: backup gate
The gallery holds face embeddings and crops of household members. esh-ml1 is
**outside vzdump**, so the gallery reaches backup only through the app's backup
CLI, which writes `gallery.db` (mode 0600) into `BACKUP_DIR` =
`/var/lib/restic/stage/augaman` (owned 10001:10001, mode 0700).
**Operator ruling: until a scheduled backup ships that file off-box AND one
restore has been verified (the restored copy reports the same identities), only
public-domain test fixtures may be enrolled. No household faces.**
Backup status (2026-09-27): **WIRED AND RESTORE-VERIFIED, so the gate is met.**
restic runs daily at 0100 PT to rest-server-ana, with a fail-closed pre-backup
hook ([`configs/restic/esh-ml1/`](../../configs/restic/esh-ml1/README.md)). The
restore was verified against augaman-dev's public-domain canary identity (1
identity, 3 samples): the copy restored from snapshot `fd3061a1` matched the
live gallery (identities, samples and embeddings, by digest), and
`integrity_check` returned ok.
```bash
docker exec augaman python -m augaman.gallery.backup --db /data/gallery.db --dest /backup/gallery.db
# one JSON line + exit 0, or "backup failed: ..." + exit 1
```
## Building
The image is built on esh-ml1 from the release tag's content. esh-ml1 has no
gitea credentials, so the source is shipped as a `git archive`:
```bash
# from nh3-dev, in a pfi/augaman checkout
git archive --format=tar vX.Y.Z | ssh esh-ml1 'sudo mkdir -p /opt/docker/src/augaman-vX.Y.Z && sudo tar -x -C /opt/docker/src/augaman-vX.Y.Z && sudo chown -R infra-ops:infra-ops /opt/docker/src'
# on esh-ml1
cd /opt/docker/src/augaman-vX.Y.Z && docker build -t augaman:X.Y.Z .
```
The build fetches the two pinned models and SHA-256-checks them; a mismatch fails
the build. v0.1.0 built in under 2 minutes cold; v0.1.1 was the first version
deployed (2026-09-26), then v0.1.2 (2026-09-27).
⚠ **Disk:** a build that has to install the third-party packages takes ~8–11 GB
transiently (image ~5 GB plus the ~3 GB uv download cache). Beszel alerts at 85%.
Before v0.1.2, every version bump re-installed them, and the v0.1.1 build hit 90%.
From v0.1.2 the packages install from a layer keyed on `uv export
--no-emit-project`, so a bump that changes no dependency reuses that layer. Keep
the layer cache (`docker builder prune --filter type=exec.cachemount` drops only
the download cache); a full `docker builder prune` forces the next build to
re-install everything.
## Startup is fail-closed on CUDA
The service refuses to serve unless CUDA really runs every convolution (profiled
warmup, ORT CPU fallback disabled). `/health` reports backend `cuda` only then.
Allow up to 180 s (`start_period`). The real GPU confirmation is the augaman
process showing up in `nvidia-smi` on the host.