feat(augaman): second, fixtures-only instance on fv-ml1 GPU 1; CPU vs GPU speed bench (v0.1.2 baseline)
Prime asked for augaman on fv-ml1's utility card, beside vllm-coder. Mirror augaman-dev's f77164f compose, which parameterises the GPU reservation (GPU_ID, default 0) and the Homepage card name (CARD_SUFFIX). esh-ml1's resolved config is unchanged: same config hash, no recreate. On fv-ml1: augaman:0.1.2 built on-box from the tag, GPU_ID=1, healthy on CUDA at 1264 MiB, and pytest -m gpu tests/vision passes 3/3 on the Blackwell. It has its own gallery and no gallery backup, so it is fixtures-only. The host's raw restic copy of /var/lib/docker/volumes is not a consistent SQLite backup. docs/pfi/augaman-speed-bench/ holds the harness (augaman-dev's recipe plus a no-face control frame and a face-count check on every response), the raw rows and the summary. Server-side, one face: - esh-ml1 GPU 144 ms - fv-ml1 GPU 75 ms - fv-ml1 CPU on 6 cores 152 ms - esh-ml1 CPU 888 ms It agrees with augaman-dev's independent esh-ml1 measurement once each harness's floor is subtracted. This is the before for v0.1.3's detector fix.
This commit is contained in:
@@ -11,9 +11,19 @@ IMAGE=augaman:0.1.2
|
||||
PORT=8040
|
||||
HOST_IP=10.0.50.80
|
||||
|
||||
# The host's card index for the GPU reservation (default 0). The container always
|
||||
# sees its one card as index 0, so AUGAMAN_CUDA_DEVICE_ID never changes.
|
||||
# esh-ml1: 0 (the only card) fv-ml1: 1 (the utility card, beside vllm-coder)
|
||||
GPU_ID=0
|
||||
# Appended to the Homepage card name so the two instances are distinguishable.
|
||||
# Note the leading space. esh-ml1 (the primary) leaves it empty.
|
||||
# fv-ml1: CARD_SUFFIX=" (fv-ml1)"
|
||||
CARD_SUFFIX=
|
||||
|
||||
# Where the backup CLI writes gallery.db. Owned 10001:10001 (the container user),
|
||||
# mode 0700. It is the restic stage dir, so the backup run that copies the gallery here
|
||||
# is the same run that ships it off-box (README "Backup").
|
||||
# mode 0700. On esh-ml1 it is the restic stage dir, so the backup run that copies the
|
||||
# gallery here is the same run that ships it off-box (README "Backup").
|
||||
# fv-ml1 (fixtures-only, no scheduled gallery backup): /opt/docker/backup/augaman.
|
||||
BACKUP_DIR=/var/lib/restic/stage/augaman
|
||||
|
||||
# >= 32 visible-ASCII characters. Source of truth is the vault:
|
||||
|
||||
@@ -13,7 +13,14 @@ first, then re-mirror it here.
|
||||
| **Token** | `secret get augaman/api-token` (vault is the source of truth) |
|
||||
| **Image** | `augaman:<version>`, built locally on esh-ml1 (below) |
|
||||
| **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) |
|
||||
| **VRAM** | ~1.5 GB by design (embed batches capped at 16) |
|
||||
| **VRAM** | ~0.5 GB on esh-ml1, ~1.3 GB on fv-ml1 (measured) |
|
||||
|
||||
**Two instances (2026-09-27).** The **primary is on esh-ml1** (`GPU_ID=0`): it holds
|
||||
the gallery and has the verified backup. A **second instance runs on fv-ml1**
|
||||
(`GPU_ID=1`, the utility card beside `vllm-coder`; `CARD_SUFFIX=" (fv-ml1)"`;
|
||||
`BACKUP_DIR=/opt/docker/backup/augaman`) at Prime's request. It has its own,
|
||||
separate gallery, and it is **fixtures-only**: no gallery backup is wired there,
|
||||
and the two galleries do not sync. Speed comparison: [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md).
|
||||
|
||||
## ⚠ Biometric data: backup gate
|
||||
|
||||
|
||||
@@ -1,7 +1,8 @@
|
||||
# augaman: the fleet's face-recognition service for Cicada (gitea pfi/augaman), on esh-ml1
|
||||
# (CT 110 on esh-pve, RTX 2000E Ada 16 GB). Enroll, recognize, verify; buffalo_l (SCRFD +
|
||||
# ArcFace w600k_r50) on ONNX Runtime CUDA. Canonical copy: this file in pfi/augaman; the
|
||||
# eshpfi stack mirrors it as stacks/augaman.
|
||||
# augaman: the fleet's face-recognition service for Cicada (gitea pfi/augaman). The primary
|
||||
# instance is on esh-ml1 (CT 110 on esh-pve, RTX 2000E Ada 16 GB), with the gallery and the
|
||||
# backup. A second, fixtures-only instance runs on fv-ml1 (GPU_ID=1, no backup). Enroll,
|
||||
# recognize, verify; buffalo_l (SCRFD + ArcFace w600k_r50) on ONNX Runtime CUDA. Canonical copy:
|
||||
# this file in pfi/augaman; the eshpfi stack mirrors it as stacks/augaman.
|
||||
#
|
||||
# ⚠ BIOMETRIC DATA. The gallery volume holds face embeddings and crops of household members.
|
||||
# - The live SQLite stays on the local named volume. Never NFS (the service refuses it).
|
||||
@@ -21,7 +22,10 @@
|
||||
# variables below carry no AUGAMAN_ prefix and are never passed through wholesale (no env_file).
|
||||
#
|
||||
# .env (tunables): IMAGE, PORT (8040), BACKUP_DIR, HOST_IP (10.0.50.80), AUGAMAN_API_TOKEN
|
||||
# (>= 32 visible-ASCII characters; the source of truth is the vault).
|
||||
# (>= 32 visible-ASCII characters; the source of truth is the vault), GPU_ID (the host's card
|
||||
# index, default 0), CARD_SUFFIX (appended to the Homepage name, e.g. " (fv-ml1)").
|
||||
# AUGAMAN_CUDA_DEVICE_ID stays "0" on every host: the reservation shows the container only the
|
||||
# card GPU_ID names, and it sees that card as index 0.
|
||||
|
||||
name: augaman
|
||||
|
||||
@@ -47,7 +51,7 @@ services:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids: ["0"]
|
||||
device_ids: ["${GPU_ID:-0}"]
|
||||
capabilities: [gpu]
|
||||
healthcheck:
|
||||
# 200 only when ready and not degraded; a 503 (degraded) fails the check.
|
||||
@@ -58,7 +62,7 @@ services:
|
||||
start_period: 180s
|
||||
labels:
|
||||
- homepage.group=AI - Eval & Retrieval
|
||||
- homepage.name=augaman — face recognition
|
||||
- homepage.name=augaman — face recognition${CARD_SUFFIX:-}
|
||||
- homepage.icon=mdi-face-recognition
|
||||
- homepage.description=Enroll, recognize, verify (buffalo_l on CUDA) for Cicada
|
||||
- homepage.href=http://${HOST_IP:-10.0.50.80}:${PORT:-8040}/health
|
||||
|
||||
Reference in New Issue
Block a user