# ana-ml2 Primary AI inference host for PFI. ## Network - **LAN IP:** 10.250.50.54 (in-band, OS-side) - **BMC (OOB):** 10.250.250.50 — Supermicro IPMI web UI at (homepage card: *PFI-ANA-ML2 BMC*) - **SSH:** standard port 22 on 10.250.50.54 ## Hardware - **Chassis:** Supermicro mid-range inferencing server (bare metal, NOT Dell / not the same box as sf-r630 / sfsrv-ana) - **CPU:** AMD EPYC 9254 24-core (96 threads) - **RAM:** 566 GB - **GPUs:** 2x NVIDIA RTX 6000 Ada Generation (46 GB VRAM each, GPU 0 and GPU 1) - **Storage:** ZFS `zroot` (434 GB root) + `tank` pool (8.6 TB at `/tank`) - **OS:** Debian 13 (trixie), kernel 6.12.x - **Docker:** 29.3.1, runtimes: runc (default), nvidia, io.containerd.runc.v2 ## Key paths | Path | Purpose | |------|---------| | `/opt/docker/compose//` | Compose files | | `/opt/docker/conf//` | Config bind mounts | | `/tank/aimodels/huggingface/` | HF cache (267 GB, pre-downloaded models) | | `/tank/aimodels/llm/` | Legacy GGUF models (790 GB, referenced by llama-swap as `/models/`) | | `/var/lib/docker/` | Docker data (on zroot) | ## Running stacks | Stack | Port | Notes | |-------|------|-------| | llama-swap | 9292 | GGUF model server via llama.cpp | | vllm-embed (Qwen3) | 8001 | OpenAI-compatible embeddings; part of the `vllm` stack (GPU 1) | | vllm-rerank (Qwen3) | 8002 | OpenAI-compatible reranker; part of the `vllm` stack (GPU 1) | | vllm-reward (Skywork) | 8003 | Skywork-Reward-V2-8B-AWQ classifier; part of the `vllm` stack (GPU 1) | | dockge | 5001 | Docker stack management UI | | dozzle-agent | 7007 | Log agent; reports to the Dozzle hub on ana-docker | | beszel-agent | 45876 | Metrics agent; reports to the Beszel hub on ana-docker | **Retired since last README update:** - `infinity` — replaced by the `vllm` stack (originally `vllm-qwen3`, renamed 2026-05-13 when the stack expanded beyond Qwen3) after the upstream Infinity image stopped shipping a `transformers` build that knew Qwen3. - `LibreChat (+ rag_api, vectordb, mongodb, meilisearch)` — removed from this host. - `searxng` — now hosted on ana-docker for the whole fleet. - Residual networks (`librechat_default`, `kokoro-tts-gpu_default`) from prior experiments are still present; safe to `docker network rm` at leisure. ## Refresh state ```bash scripts/refresh-server-info.sh ana-ml2 ``` Latest snapshot: `system-details.txt` (regenerate as needed). ## GPU allocation policy By default, no container is pinned. For predictable performance when multiple GPU workloads run concurrently: - **GPU 0:** heavy LLM (llama-swap big models). - **GPU 1:** light services (the three `vllm` services share this GPU via `--gpu-memory-utilization`). Use `deploy.resources.reservations.devices[].device_ids: [""]` in compose to pin.