Files
esh-pfi-infrastructure/services/intern-decision-serve/acceptance/gpu3-2026-09-30/poll.sh
T
vh a262477a61 feat(intern-decision): stack, DNS and GPU 3 acceptance for the SemIf replacement
stacks/intern-decision: compose (GPU 1, :8033, hard VRAM cap as the single .env knob,
healthcheck, Homepage group 'AI - Eval & Retrieval'), .env.example and README.
dns: intern-decision.fv.internal -> fv-ml1 (synced to ana/esh/nh3).
acceptance on fv-ml1 GPU 3, 3 fresh processes: bit-identical to the Jev bench's native rows
(pooled 240/259, Wyrd 79/84, 0/560 flips, Δp 0), negative control 10/122/14, 0 flips across
restarts; largest accepted request 200 at a 10,134 MiB card peak under a 9.25 GiB cap; 503 and
recovery proven at a tight cap. GPU 1 deploy held: nvidia-smi Free on GPU 1 is 15,442 MiB.
2026-09-30 09:38:00 -07:00

11 lines
657 B
Bash
Executable File

#!/bin/bash
# poll.sh <container> <out.csv>: unix time, that container's process GPU memory (MiB, nvidia-smi per process), card total
pid=$(docker inspect -f '{{.State.Pid}}' "$1")
gpu=$(nvidia-smi --query-compute-apps=pid,gpu_uuid --format=csv,noheader | awk -F', ' -v p="$pid" '$1==p{print $2}')
echo "# pid=$pid gpu=$gpu" > "$2"
while [ -d /proc/$pid ]; do
procmem=$(nvidia-smi --query-compute-apps=pid,used_memory --format=csv,noheader,nounits | awk -F', ' -v p="$pid" '$1==p{print $2}')
card=$(nvidia-smi --id="$gpu" --query-gpu=memory.used --format=csv,noheader,nounits)
printf '%s,%s,%s\n' "$(date +%s.%N)" "${procmem:-0}" "$card" >> "$2"
done