fv-ml1: complete the cutover — rename, renumber, DNS, and the LiteLLM repoint

The box is physically at Fountain Valley, renamed, renumbered onto 10.251/16,
and serving inference again. This lands the repo half of that.

Host: hostname ana-ml2 -> fv-ml1, pinned to 10.251.50.54 by a dnsmasq
reservation so the address the runbook, DNS and LiteLLM all assume is the
address it actually has. Its headscale node is renamed too.

The sweep ran from scripts/fv-ml1-rename-sweep.sh, whose allowlist is the
reason this diff touches current-state files and not the record. Dated
persistent-memory entries, archival-memory and incident notes still say
ana-ml2 in 31 and 62 places respectively, because that is what the box was
when those things happened. Rewriting them would make the history lie.

LiteLLM was the load-bearing piece and needed more than the api_base sed the
runbook describes. Twenty api_base entries repointed, but a grep-and-verify
pass also caught a LIVE pass_through_endpoints target for the scalar-judge
reward route still on the old address -- an api_base-only substitution would
have left it dead. Four prose references describing current state were
repointed as well; one historical note recording where a hand-test was run
is deliberately left pointing at 10.250.50.54.

Two facts in the server tables were wrong and are corrected here. The site is
Fountain Valley, not Anaheim. And the box has FOUR RTX PRO 6000 Blackwell
Max-Q, not two -- verified by nvidia-smi -L and independently by PCI
enumeration of four GB202GL devices. That is 391 GB of VRAM rather than 196,
which changes what fits on it.

DNS: fv-ml1, fv-ml1-bmc and fv-gw added under the fv site via the piggyback
approach, scriberr re-homed, and the ana-ml2 records removed. Applied to all
three resolvers. The BMC record carries a warning that its 802.1q VLAN tag
must stay disabled -- it shipped tagging VLAN 250 into an untagged port,
which made it invisible to every network-side diagnostic and is the reason
it appeared dead through several cable changes.

Verified end to end: summarizer and sec both answer through the Anaheim
gateway across the mesh to FV seats on different ports.
This commit is contained in:
vh
2026-09-12 22:00:50 -07:00
parent fde7834635
commit 91bda3c480
31 changed files with 115 additions and 112 deletions
+182
View File
@@ -0,0 +1,182 @@
# ana-ml2 — OPEN the ERP/RP tune window: clear GPU0 completely.
#
# stop mog-sec (GPU1) -> move gen GPU0 -> GPU1 -> GPU0 empty for training
#
# Operator call 2026-08-24: rather than train beside `gen`, move `gen` off GPU0
# entirely and stand `sec` down for the night. Training then gets a whole card
# (95.60 GiB) instead of a shared one, and the fleet's general seat never goes
# dark beyond its own restart.
#
# ⚠⚠ ORDER IS LOAD-BEARING — DO NOT REORDER THE STEPS.
# `gen` runs at --gpu-memory-utilization 0.43, which vLLM reads as a fraction of
# TOTAL card memory: 0.43 x 97,887 MiB = 42,091 MiB that must be FREE at startup
# or the engine refuses to boot. GPU1 has only 19,446 MiB free while mog-sec is
# up. Recreating `gen` onto GPU1 first would take the fleet's main seat down and
# leave it down. mog-sec stops FIRST, and step 3 hard-gates on the freed memory
# before `gen` is touched at all.
#
# ⚠ `stop`, never `down`. `down` removes the container; `stop` leaves it in
# place so the close playbook can `start` it. Both seats are `restart:
# unless-stopped`, which does NOT resurrect a deliberately-stopped container.
#
# ⚠ device_ids vs nvidia-smi ordering was VERIFIED on this host, not assumed:
# gen (GEN_GPU_ID=0) reports under the GPU nvidia-smi indexes 0, mog-sec
# (MOG_GPU_ID=1) under index 1. They agree here. (They do NOT on irv-ml1 —
# never carry that assumption between boxes.)
#
# Restore with: playbooks/ana-ml2-training-window-close.yaml
#
# scripts/elway ana-ml2 --playbook playbooks/ana-ml2-training-window-open.yaml
vars:
gen_dir: /opt/docker/compose/gen-seat
mog_dir: /opt/docker/compose/mog-sec
gen_port: "8015"
# gen's startup requirement: 0.43 x 97,887 MiB, rounded up. If GPU1 has less
# than this free, gen will not boot and the window must not proceed.
gen_required_free_mib: "42100"
steps:
- name: "PREFLIGHT — GPU0 holds vllm-gen and nothing else unexpected"
sudo: true
shell: |
set -e
procs=$(nvidia-smi --query-compute-apps=gpu_uuid,pid --format=csv,noheader | wc -l)
gpu0_uuid=$(nvidia-smi --query-gpu=uuid --format=csv,noheader -i 0)
gpu0_procs=$(nvidia-smi --query-compute-apps=gpu_uuid,pid --format=csv,noheader | grep -c "$gpu0_uuid" || true)
echo "GPU0 compute procs: $gpu0_procs (total on box: $procs)"
test "$gpu0_procs" -le 1
changed_when: "false"
- name: "PREFLIGHT — record gen's current container id (proves the recreate later)"
sudo: true
shell: docker inspect vllm-gen --format '{{.Id}}' | tee /tmp/gen-container-id-before.txt
changed_when: "false"
- name: "Stop mog-sec (the `sec` / `sec-reasoning` seat) — frees ~55.3 GiB on GPU1"
sudo: true
shell: docker compose --project-directory {{ mog_dir }} stop vllm-mog-sec
# Skip if already stopped, so the playbook is rerunnable.
when: "docker inspect -f '{{.State.Running}}' vllm-mog-sec 2>/dev/null | grep -q true"
- name: "Wait for GPU1 memory to actually release (teardown is not instant)"
sudo: true
shell: |
for i in $(seq 1 60); do
free=$(nvidia-smi --query-gpu=memory.free --format=csv,noheader,nounits -i 1)
if [ "$free" -ge {{ gen_required_free_mib }} ]; then
echo "GPU1 free: ${free} MiB"; exit 0
fi
sleep 2
done
echo "TIMEOUT: GPU1 free is ${free} MiB, need >= {{ gen_required_free_mib }}"; exit 1
changed_when: "false"
- name: "HARD GATE — GPU1 has room for gen's 0.43 budget before we touch gen"
sudo: true
shell: |
free=$(nvidia-smi --query-gpu=memory.free --format=csv,noheader,nounits -i 1)
echo "GPU1 free ${free} MiB vs required {{ gen_required_free_mib }} MiB"
test "$free" -ge {{ gen_required_free_mib }}
changed_when: "false"
- name: "Point gen at GPU1 in its .env (GEN_GPU_ID 0 -> 1)"
sudo: true
# ⚠ `sudo` INSIDE the when: expression. A step's `sudo: true` covers the
# shell, NOT its when/creates/changed_when guards — those run as the login
# user. The .env is root-only 0600, so an unsudo'd grep exits 2
# (permission denied), which is not 0, so the step SILENTLY SKIPS and the
# flip never happens. Caught 2026-08-24 by the effective-value assert below.
shell: sed -i 's/^GEN_GPU_ID=0$/GEN_GPU_ID=1/' {{ gen_dir }}/.env
when: "sudo grep -qx 'GEN_GPU_ID=0' {{ gen_dir }}/.env"
- name: "Assert the EFFECTIVE device id, not the .env line"
sudo: true
# grep on the .env proves a substring is present; only `compose config`
# proves what the container will actually be created with.
# ⚠ Parse the JSON, do not regex the YAML. The first version of this grepped
# for -\s*'?1'? and failed against compose's DOUBLE-quoted `- "1"` — an
# assert that fails for the wrong reason is worse than no assert.
shell: |
docker compose --project-directory {{ gen_dir }} config --format json \
| jq -e '.services["vllm-gen"].deploy.resources.reservations.devices[0].device_ids == ["1"]'
changed_when: "false"
- name: "Recreate gen onto GPU1 (a device change needs up -d, not restart)"
sudo: true
shell: docker compose --project-directory {{ gen_dir }} up -d vllm-gen
- name: "Wait for gen to serve /health (cold start: weights + CUDA graphs + MTP)"
sudo: true
shell: |
for i in $(seq 1 180); do
if curl -sf -o /dev/null http://127.0.0.1:{{ gen_port }}/health; then
echo "gen healthy after $((i*5))s"; exit 0
fi
sleep 5
done
echo "TIMEOUT: gen did not become healthy in 900s"; exit 1
changed_when: "false"
verify:
- name: "gen was genuinely RECREATED (container id changed)"
sudo: true
shell: |
before=$(cat /tmp/gen-container-id-before.txt)
after=$(docker inspect vllm-gen --format '{{.Id}}')
echo "before=${before:0:12} after=${after:0:12}"
test "$before" != "$after"
changed_when: "false"
- name: "gen's process is resident on GPU1"
sudo: true
# ⚠ Match by CGROUP, not by `.State.Pid`. vLLM V1 runs EngineCore as a CHILD
# of the container's pid 1, and it is the child that holds the GPU memory —
# nvidia-smi never reports `.State.Pid`, so comparing against it always fails.
shell: |
cid=$(docker inspect vllm-gen --format '{{.Id}}')
gpu1_uuid=$(nvidia-smi --query-gpu=uuid --format=csv,noheader -i 1)
found=0
for p in $(nvidia-smi --query-compute-apps=gpu_uuid,pid,used_memory --format=csv,noheader \
| grep "$gpu1_uuid" | cut -d, -f2 | tr -d ' '); do
if grep -q "$cid" /proc/$p/cgroup 2>/dev/null; then
echo "gen pid $p resident on GPU1: $(nvidia-smi --query-compute-apps=pid,used_memory --format=csv,noheader | grep "^$p,")"
found=1
fi
done
test "$found" -eq 1
changed_when: "false"
- name: "gen answers a real completion, not just /health"
sudo: true
shell: |
. {{ gen_dir }}/.env
curl -sf -m 120 http://127.0.0.1:{{ gen_port }}/v1/chat/completions \
-H "Authorization: Bearer ${API_KEY}" -H 'Content-Type: application/json' \
-d '{"model":"'"${GEN_SERVED_NAME}"'","messages":[{"role":"user","content":"reply with the single word: ok"}],"max_tokens":16}' \
| grep -q '"content"'
changed_when: "false"
- name: "GPU0 IS EMPTY — zero compute processes"
sudo: true
shell: |
gpu0_uuid=$(nvidia-smi --query-gpu=uuid --format=csv,noheader -i 0)
n=$(nvidia-smi --query-compute-apps=gpu_uuid,pid --format=csv,noheader | grep -c "$gpu0_uuid" || true)
free=$(nvidia-smi --query-gpu=memory.free --format=csv,noheader,nounits -i 0)
echo "GPU0 compute procs=${n} free=${free} MiB"
test "$n" -eq 0 && test "$free" -ge 95000
changed_when: "false"
- name: "mog-sec is stopped (not removed — close depends on `start` working)"
sudo: true
shell: |
docker inspect -f '{{.State.Status}}' vllm-mog-sec | tee /dev/stderr | grep -qx exited
changed_when: "false"
- name: "GPU1 still has headroom for Scriberr's on-demand load"
sudo: true
shell: |
free=$(nvidia-smi --query-gpu=memory.free --format=csv,noheader,nounits -i 1)
echo "GPU1 free after gen landed: ${free} MiB"
test "$free" -ge 12000
changed_when: "false"