chore(augaman): remove the fv-ml1 instance (Prime)

Prime removed the second instance after the v0.1.3 bench. esh-ml1 handles a face
in ~48 ms, sits in the house next to the cameras, and holds the verified
backup. fv-ml1's gallery was empty (0 identities). The container, gallery volume,
image, compose dir (with its .env), backup dir and build sources are removed from
fv-ml1. GPU_ID / CARD_SUFFIX stay in the compose for any future second host.
This commit is contained in:
vh
2026-09-27 01:31:25 -07:00
parent 0118e57c83
commit ad484c3c99
4 changed files with 15 additions and 22 deletions
+3 -3
View File
@@ -155,12 +155,12 @@ _As of 2026-09-26 ~1620 PT._
survived the recreate and was then deleted, so the gallery is empty and ready for real survived the recreate and was then deleted, so the gallery is empty and ready for real
enrollments. `/recognize` p50 186 ms (1080p, one face). The detector-latency follow-up enrollments. `/recognize` p50 186 ms (1080p, one face). The detector-latency follow-up
is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops. is theirs. Their next release pins the container gid to 10001; nothing is owed by infra-ops.
- **Second instance on fv-ml1 GPU 1 (Prime, 2026-09-27)**, `GPU_ID=1`, **fixtures-only, no - **The fv-ml1 instance was REMOVED by Prime on 2026-09-27 (0125 PT)** after the bench. esh-ml1 is the only
gallery backup**, separate gallery. Its role (bench-only vs future primary) is Prime's call. instance: in the house, backed up, 48 ms per face on v0.1.3.
- **v0.1.3 live on BOTH hosts (2026-09-27 ~0105 PT).** Speed bench (`docs/pfi/augaman-speed-bench/`), - **v0.1.3 live on BOTH hosts (2026-09-27 ~0105 PT).** Speed bench (`docs/pfi/augaman-speed-bench/`),
server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode server-side one face, v0.1.2 → v0.1.3: esh GPU 144 → 48 ms, fv GPU 75 → 27 ms. CPU mode
REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to REGRESSED (fv CPU6 152 → 205 ms; suspected ORT thread-pool spinning) and is reported to
augaman-dev; neither deployment uses CPU mode. The fv-ml1 instance's role is still Prime's call. augaman-dev; the deployment does not use CPU mode.
### esh-ml1 ### esh-ml1
+3 -8
View File
@@ -180,14 +180,9 @@ embed/rerank/reward trio. GPUs are pinned per container via
**GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):** **GPU 1 — light / eval / retrieval + char-RP GGUF (~91/98 GB, on-demand):**
> ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`, > ⚠ **This table is stale (checked 2026-09-27).** Live GPU 1 residents were `scriberr`,
> `vllm-coder`, `vllm-erp-seat`, `vllm-meromero-rp` and now **`augaman`** (below). Read the > `vllm-coder`, `vllm-erp-seat` and `vllm-meromero-rp`. Read the host
> host (`docker inspect … DeviceRequests`), not this table. > (`docker inspect … DeviceRequests`), not this table. (A fixtures-only `augaman` instance ran
> > here for about an hour on 2026-09-27 for a speed bench, and was then removed on Prime's call.)
> **`augaman` :8040 (since 2026-09-27, Prime):** the second instance of the face-recognition service
> (`stacks/augaman`, `GPU_ID=1`), ~1.3 GB. **Fixtures-only: it has no gallery backup wired.**
> The primary instance, which holds the gallery and its backup, is on esh-ml1. Note that this host's restic copies
> `/var/lib/docker/volumes` raw, and that includes `augaman_gallery`: a live SQLite file, so the copy is
> not guaranteed consistent. That is acceptable for fixtures and not for real faces.
| Container | Port | Served model | Quant | Ctx | | Container | Port | Served model | Quant | Ctx |
|-----------|------|--------------|-------|-----| |-----------|------|--------------|-------|-----|
+2 -4
View File
@@ -13,17 +13,15 @@ HOST_IP=10.0.50.80
# The host's card index for the GPU reservation (default 0). The container always # The host's card index for the GPU reservation (default 0). The container always
# sees its one card as index 0, so AUGAMAN_CUDA_DEVICE_ID never changes. # sees its one card as index 0, so AUGAMAN_CUDA_DEVICE_ID never changes.
# esh-ml1: 0 (the only card) fv-ml1: 1 (the utility card, beside vllm-coder) # esh-ml1: 0 (the only card). fv-ml1 used 1 while its instance existed (removed 2026-09-27).
GPU_ID=0 GPU_ID=0
# Appended to the Homepage card name so the two instances are distinguishable. # Appended to the Homepage card name so the two instances are distinguishable.
# Note the leading space. esh-ml1 (the primary) leaves it empty. # Note the leading space, e.g. CARD_SUFFIX=" (fv-ml1)". esh-ml1 leaves it empty.
# fv-ml1: CARD_SUFFIX=" (fv-ml1)"
CARD_SUFFIX= CARD_SUFFIX=
# Where the backup CLI writes gallery.db. Owned 10001:10001 (the container user), # Where the backup CLI writes gallery.db. Owned 10001:10001 (the container user),
# mode 0700. On esh-ml1 it is the restic stage dir, so the backup run that copies the # mode 0700. On esh-ml1 it is the restic stage dir, so the backup run that copies the
# gallery here is the same run that ships it off-box (README "Backup"). # gallery here is the same run that ships it off-box (README "Backup").
# fv-ml1 (fixtures-only, no scheduled gallery backup): /opt/docker/backup/augaman.
BACKUP_DIR=/var/lib/restic/stage/augaman BACKUP_DIR=/var/lib/restic/stage/augaman
# >= 32 visible-ASCII characters. Source of truth is the vault: # >= 32 visible-ASCII characters. Source of truth is the vault:
+7 -7
View File
@@ -13,14 +13,14 @@ first, then re-mirror it here.
| **Token** | `secret get augaman/api-token` (vault is the source of truth) | | **Token** | `secret get augaman/api-token` (vault is the source of truth) |
| **Image** | `augaman:<version>`, built locally on esh-ml1 (below) | | **Image** | `augaman:<version>`, built locally on esh-ml1 (below) |
| **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) | | **State** | named volume `augaman_gallery` (SQLite, local disk; the app refuses NFS) |
| **VRAM** | ~0.5 GB on esh-ml1, ~1.3 GB on fv-ml1 (measured) | | **VRAM** | ~0.74 GB on v0.1.3 (measured after the bench; 0.51 GB on v0.1.2) |
**Two instances (2026-09-27).** The **primary is on esh-ml1** (`GPU_ID=0`): it holds **One instance: esh-ml1.** A second, fixtures-only instance ran on fv-ml1's GPU 1
the gallery and has the verified backup. A **second instance runs on fv-ml1** (`GPU_ID=1`) from 2026-09-27 0017 to 0125 PT, for the speed comparison in
(`GPU_ID=1`, the utility card beside `vllm-coder`; `CARD_SUFFIX=" (fv-ml1)"`; [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md).
`BACKUP_DIR=/opt/docker/backup/augaman`) at Prime's request. It has its own, **Prime removed it** after v0.1.3 put esh-ml1 at ~48 ms per face: it is in the house
separate gallery, and it is **fixtures-only**: no gallery backup is wired there, next to the cameras and it holds the verified backup. `GPU_ID` / `CARD_SUFFIX` stay in
and the two galleries do not sync. Speed comparison: [`docs/pfi/augaman-speed-bench/`](../../docs/pfi/augaman-speed-bench/README.md). the compose for any future second host.
## ⚠ Biometric data: backup gate ## ⚠ Biometric data: backup gate