355a2407a2
ana-ml2 was upgraded 2026-06 from dual RTX 6000 Ada (46GB, cc 8.9) to
dual RTX PRO 6000 Blackwell Max-Q (96GB, cc 12.0 / sm_120). Update the
stale hardware facts across the workspace:
- CLAUDE.md servers table row
- servers/ana-ml2/README.md hardware spec (+ refreshed system-details.txt)
- stacks/vllm compose + .env.example FP8/KV comments (Ada cc 8.9 -> Blackwell cc 12.0)
- stacks/llama-swap config VRAM-budget comment (48GB -> 96GB, GPU-0 pin)
Also corrects the adjacent stale 'Phi-4-mini' comment in the granite
service block (the service has been Granite 4.1 8B since 34a43a0).
Doc/comment-only; no runtime change.
1044 lines
42 KiB
Plaintext
1044 lines
42 KiB
Plaintext
|
|
===== HOST =====
|
|
|
|
Hostname: ana-ml2
|
|
Date: 2026-06-13T13:35:42-07:00
|
|
Uptime: up 1 day, 6 minutes
|
|
OS: Debian GNU/Linux 13 (trixie)
|
|
Kernel: 6.12.74+deb13+1-amd64
|
|
Arch: x86_64
|
|
|
|
===== HARDWARE =====
|
|
|
|
CPU cores: 96
|
|
CPU model: AMD EPYC 9254 24-Core Processor
|
|
MemTotal: 566.6 GB
|
|
MemAvailable: 484.5 GB
|
|
|
|
===== GPUS =====
|
|
|
|
index, name, memory.total [MiB], memory.free [MiB], driver_version
|
|
0, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 97247 MiB, 580.65.06
|
|
1, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 3692 MiB, 580.65.06
|
|
|
|
===== FILESYSTEMS (df) =====
|
|
|
|
Filesystem Size Used Avail Use% Mounted on
|
|
zroot/ROOT/debian 394G 212G 182G 54% /
|
|
efivarfs 128K 67K 57K 55% /sys/firmware/efi/efivars
|
|
/dev/sdb1 511M 92M 420M 18% /boot/efi
|
|
tank 8.6T 1.6T 7.1T 18% /tank
|
|
zroot/home 244G 62G 182G 26% /home
|
|
|
|
===== PERSISTENT MOUNTS (/etc/fstab, non-comment) =====
|
|
|
|
UUID="3D9B-8E0C" /boot/efi vfat defaults 0 0
|
|
|
|
===== TARGETED DATA PATHS =====
|
|
|
|
/tank (total: 1.6T)
|
|
total 15
|
|
drwxrwxrwx 6 root root 6 2026-04-11 23:17 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
drwxrwxr-x 9 llmuser llm 10 2026-06-12 18:01 aimodels
|
|
drwxrwxr-x 3 llmuser llm 3 2025-09-08 13:21 comfy
|
|
drwxrwxr-x 4 llmuser llmuser 4 2026-04-11 23:17 kokoro
|
|
drwxrwxr-x 5 lkraven lkraven 6 2025-09-10 18:35 vibevoice
|
|
|
|
/opt (total: 16G)
|
|
total 79
|
|
drwxrwxrwx 13 root root 13 2026-04-17 23:00 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
drwx--x--x 4 root root 4 2025-09-02 12:53 containerd
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 docker
|
|
drwxrwxr-x 8 llmuser llm 17 2026-04-17 17:50 heretic
|
|
drwxrwxr-x 15 llmuser llmuser 36 2026-04-11 23:24 Kokoro-FastAPI
|
|
drwxrwxr-x 20 llmuser llmuser 39 2025-10-09 21:42 LibreChat
|
|
drwxrwxr-x 26 llmuser llmuser 56 2025-09-04 21:01 llama.cpp
|
|
drwxrwxr-x 13 llmuser llmuser 25 2025-09-02 12:50 llama-swap
|
|
drwxrwxr-x 3 llmuser llmuser 11 2026-04-18 15:48 llmcompressor
|
|
drwxr-xr-x 4 root root 4 2025-08-30 22:54 nvidia
|
|
drwxrwxr-x 6 llmuser llmuser 15 2025-09-24 09:53 parakeet-tdt-0.6b-v2-fastapi
|
|
drwxrwxr-x 2 llmuser llmuser 6 2025-09-05 10:03 uv
|
|
|
|
/opt/docker (total: 110M)
|
|
total 18
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 .
|
|
drwxrwxrwx 13 root root 13 2026-04-17 23:00 ..
|
|
drwxrwxr-x 12 lkraven lkraven 12 2026-06-13 02:32 compose
|
|
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 conf
|
|
|
|
/opt/docker/compose (total: 110M)
|
|
total 30
|
|
drwxrwxr-x 12 lkraven lkraven 12 2026-06-13 02:32 .
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 18:47 beszel-agent-ana
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-05 20:36 comfyui
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:47 dockge
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:44 dozzle-agent-ana
|
|
drwxrwxr-x 3 lkraven lkraven 4 2026-04-11 23:17 kokoro
|
|
drwxr-xr-x 2 lkraven lkraven 6 2026-06-12 15:17 llama-swap
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-24 12:32 parakeet
|
|
drwxrwxr-x 2 lkraven lkraven 4 2026-06-13 08:34 qwen35-vl
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-10 18:03 vibevoice
|
|
drwxr-xr-x 2 lkraven lkraven 9 2026-06-13 08:42 vllm
|
|
|
|
/opt/docker/conf (total: 14K)
|
|
total 2
|
|
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 .
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
|
|
drwxr-xr-x 2 lkraven lkraven 3 2026-06-05 00:22 llama-swap
|
|
drwxr-xr-x 2 lkraven lkraven 2 2026-06-04 00:37 vllm
|
|
|
|
/var/lib/docker (total: 8.5K)
|
|
|
|
/srv (total: 512)
|
|
total 9
|
|
drwxr-xr-x 2 root root 2 2025-08-30 18:46 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
|
|
|
|
===== DOCKER =====
|
|
|
|
Server: 29.3.1 Client: 29.3.1
|
|
|
|
----- docker info -----
|
|
Containers: 9 (running 9, paused 0, stopped 0)
|
|
Images: 42
|
|
Runtimes: map[io.containerd.runc.v2:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} nvidia:{{nvidia-container-runtime [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} runc:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]}]
|
|
Default runtime: runc
|
|
Storage driver: overlay2
|
|
Root dir: /var/lib/docker
|
|
Server version: 29.3.1
|
|
|
|
----- running containers -----
|
|
NAMES IMAGE STATUS PORTS
|
|
vllm-qwen35 vllm/vllm-openai Up About an hour (healthy) 0.0.0.0:8007->8000/tcp, [::]:8007->8000/tcp
|
|
vllm-granite vllm/vllm-openai:latest Up About an hour (healthy) 0.0.0.0:8004->8000/tcp, [::]:8004->8000/tcp
|
|
vllm-rerank vllm/vllm-openai:latest Up 13 hours (healthy) 0.0.0.0:8002->8000/tcp, [::]:8002->8000/tcp
|
|
vllm-embed vllm/vllm-openai:latest Up 13 hours (healthy) 0.0.0.0:8001->8000/tcp, [::]:8001->8000/tcp
|
|
vllm-reward vllm/vllm-openai:latest Up 13 hours (healthy) 0.0.0.0:8003->8000/tcp, [::]:8003->8000/tcp
|
|
llama-swap ghcr.io/mostlygeek/llama-swap:cuda Up 13 hours (healthy) 0.0.0.0:9292->8080/tcp, [::]:9292->8080/tcp
|
|
dockge louislam/dockge:latest Up 13 hours (healthy) 0.0.0.0:5001->5001/tcp, [::]:5001->5001/tcp
|
|
dozzle-agent amir20/dozzle:latest Up 13 hours 0.0.0.0:7007->7007/tcp, 8080/tcp
|
|
beszel-agent henrygd/beszel-agent:latest Up 13 hours (healthy)
|
|
|
|
----- all containers -----
|
|
NAMES IMAGE STATUS
|
|
vllm-qwen35 vllm/vllm-openai Up About an hour (healthy)
|
|
vllm-granite vllm/vllm-openai:latest Up About an hour (healthy)
|
|
vllm-rerank vllm/vllm-openai:latest Up 13 hours (healthy)
|
|
vllm-embed vllm/vllm-openai:latest Up 13 hours (healthy)
|
|
vllm-reward vllm/vllm-openai:latest Up 13 hours (healthy)
|
|
llama-swap ghcr.io/mostlygeek/llama-swap:cuda Up 13 hours (healthy)
|
|
dockge louislam/dockge:latest Up 13 hours (healthy)
|
|
dozzle-agent amir20/dozzle:latest Up 13 hours
|
|
beszel-agent henrygd/beszel-agent:latest Up 13 hours (healthy)
|
|
|
|
----- networks -----
|
|
NAME DRIVER SCOPE
|
|
bridge bridge local
|
|
host host local
|
|
kokoro-tts-gpu_default bridge local
|
|
librechat_default bridge local
|
|
llama-swap_default bridge local
|
|
none null local
|
|
traefik-net bridge local
|
|
|
|
----- networks (external, non-default — worth knowing for compose external: true) -----
|
|
kokoro-tts-gpu_default
|
|
librechat_default
|
|
llama-swap_default
|
|
traefik-net
|
|
|
|
----- named volumes -----
|
|
VOLUME NAME DRIVER
|
|
beszel-agent-ana_beszel_agent_data local
|
|
dockge_dockge_data local
|
|
dozzle-agent-ana_dozzle_agent_data local
|
|
parakeet_parakeet_cache local
|
|
searxng_searxng-data local
|
|
|
|
----- compose projects currently running -----
|
|
beszel-agent-ana
|
|
dockge
|
|
dozzle-agent-ana
|
|
llama-swap
|
|
qwen35-vl
|
|
vllm
|
|
|
|
===== COMPOSE FILES (/opt/docker/compose/) =====
|
|
|
|
|
|
>>> /opt/docker/compose/beszel-agent-ana/compose.yaml
|
|
# Beszel — lightweight server/container monitoring.
|
|
#
|
|
# Hub: single web UI with the SQLite store. Agents: per-host metric collectors
|
|
# that the hub pulls from over SSH.
|
|
#
|
|
# Multi-host layout via compose profiles:
|
|
# COMPOSE_PROFILES=hub → hub only (ana-docker)
|
|
# COMPOSE_PROFILES=hub,agent → hub + local agent on the same host
|
|
# COMPOSE_PROFILES=agent → agent only (ana-ml2, nh3-docker,
|
|
# esh-docker-vm, vm-esh-nas)
|
|
#
|
|
# The agent uses network_mode: host so it sees real host CPU/mem/net/disk
|
|
# counters rather than container-scoped ones — that's why it can't share
|
|
# the tnet network with the hub.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
beszel:
|
|
image: henrygd/beszel:${BESZEL_VERSION}
|
|
container_name: beszel
|
|
profiles: [hub]
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${BESZEL_PORT}:8090"
|
|
volumes:
|
|
- beszel_data:/beszel_data
|
|
healthcheck:
|
|
# Hub image is distroless — no wget/curl. Use the bundled `/beszel`
|
|
# binary's built-in health subcommand (https://beszel.dev/guide/healthchecks).
|
|
test: ["CMD", "/beszel", "health", "--url", "http://localhost:8090"]
|
|
interval: 120s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Monitoring
|
|
- homepage.name=Beszel
|
|
- homepage.icon=mdi-chart-line
|
|
- homepage.description=Server + container monitoring
|
|
- homepage.href=http://10.250.50.70:${BESZEL_PORT}
|
|
|
|
beszel-agent:
|
|
image: henrygd/beszel-agent:${BESZEL_VERSION}
|
|
container_name: beszel-agent
|
|
profiles: [agent]
|
|
restart: unless-stopped
|
|
network_mode: host
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- beszel_agent_data:/var/lib/beszel-agent
|
|
environment:
|
|
# Agent auth has two modes (v0.13+ supports both side-by-side):
|
|
# - KEY-mode: agent listens, hub connects inbound over SSH using KEY.
|
|
# Requires BESZEL_HUB_KEY in .env.
|
|
# - Token-mode: agent initiates an outbound connection to HUB_URL
|
|
# using TOKEN. Easier through NAT. Requires HUB_URL + BESZEL_TOKEN.
|
|
# Leave unused ones empty ("") in .env; both can be set simultaneously.
|
|
- PORT=${BESZEL_AGENT_PORT:-45876}
|
|
- KEY=${BESZEL_HUB_KEY:-}
|
|
- HUB_URL=${HUB_URL:-}
|
|
- TOKEN=${BESZEL_TOKEN:-}
|
|
- EXTRA_FILESYSTEMS=${BESZEL_EXTRA_FS:-}
|
|
healthcheck:
|
|
# Agent image ships the `/agent` binary with a `health` subcommand.
|
|
# Verifies the agent process is up — not that the hub can reach it.
|
|
test: ["CMD", "/agent", "health"]
|
|
interval: 120s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
|
|
volumes:
|
|
beszel_data:
|
|
beszel_agent_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/comfyui/compose.yaml
|
|
services:
|
|
comfyui:
|
|
runtime: nvidia
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
- compute
|
|
- utility
|
|
ports:
|
|
- 8188:8188
|
|
image: mmartial/comfyui-nvidia-docker:ubuntu24_cuda13.0-latest
|
|
networks:
|
|
- tnet
|
|
volumes:
|
|
- /tank/comfy/run:/comfy/mnt
|
|
- /tank/aimodels/img/comfy:/basedir
|
|
#user: 1001:1002
|
|
environment:
|
|
- WANTED_UID=1001
|
|
- WANTED_GID=1002
|
|
- BASE_DIRECTORY=/basedir
|
|
- SECURITY_LEVEL=weak
|
|
- NVIDIA_VISIBLE_DEVICES=all
|
|
- NVIDIA_DRIVER_CAPABILITIES=all
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=ComfyUI
|
|
- homepage.icon=mdi-panorama-variant-outline
|
|
- homepage.description=ComfyUI Image Gen (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:8188
|
|
restart: unless-stopped
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/dockge/compose.yaml
|
|
# Dockge — per-host Docker Compose UI (https://dockge.kuma.pet/).
|
|
#
|
|
# One instance runs on every Docker host so the compose dir is manageable
|
|
# from a browser. Each host sets DOCKGE_HOST_LABEL + DOCKGE_HOST_IP in its
|
|
# .env so the homepage card points at the right place.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
dockge:
|
|
image: louislam/dockge:${DOCKGE_VERSION:-latest}
|
|
container_name: dockge
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${DOCKGE_PORT:-5001}:5001"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
- dockge_data:/app/data
|
|
- /opt/docker:/opt/docker
|
|
environment:
|
|
- DOCKGE_STACKS_DIR=/opt/docker/compose
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Service Networking
|
|
- homepage.name=Dockge (${DOCKGE_HOST_LABEL})
|
|
- homepage.icon=sh-dockge.png
|
|
- homepage.description=Compose UI on ${DOCKGE_HOST_LABEL}
|
|
- homepage.href=http://${DOCKGE_HOST_IP}:${DOCKGE_PORT:-5001}
|
|
|
|
volumes:
|
|
dockge_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/dozzle-agent-ana/compose.yaml
|
|
# Dozzle — container log viewer.
|
|
#
|
|
# Multi-host layout via compose profiles:
|
|
# COMPOSE_PROFILES=hub → runs the web UI (deploy on ana-docker)
|
|
# COMPOSE_PROFILES=agent → runs the remote agent (deploy on ana-ml2)
|
|
#
|
|
# Same compose.yaml on both servers; per-host `.env` picks the profile.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
dozzle:
|
|
image: amir20/dozzle:${DOZZLE_VERSION}
|
|
container_name: dozzle
|
|
profiles: [hub]
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${DOZZLE_PORT}:8080"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- dozzle_data:/data
|
|
environment:
|
|
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
|
|
- DOZZLE_REMOTE_AGENT=${DOZZLE_REMOTE_AGENT:-}
|
|
- DOZZLE_AUTH_PROVIDER=${DOZZLE_AUTH_PROVIDER:-none}
|
|
- DOZZLE_USERNAME=${DOZZLE_USERNAME:-}
|
|
- DOZZLE_PASSWORD=${DOZZLE_PASSWORD:-}
|
|
healthcheck:
|
|
test: ["CMD", "/dozzle", "healthcheck"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Monitoring
|
|
- homepage.name=Dozzle
|
|
- homepage.icon=mdi-text-box-search
|
|
- homepage.description=Container logs (ana-docker + ana-ml2)
|
|
- homepage.href=http://10.250.50.70:${DOZZLE_PORT}
|
|
|
|
dozzle-agent:
|
|
image: amir20/dozzle:${DOZZLE_VERSION}
|
|
container_name: dozzle-agent
|
|
profiles: [agent]
|
|
restart: unless-stopped
|
|
command: agent
|
|
ports:
|
|
- "${DOZZLE_AGENT_BIND:-0.0.0.0}:${DOZZLE_AGENT_PORT}:7007"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- dozzle_agent_data:/data
|
|
environment:
|
|
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
|
|
networks:
|
|
- tnet
|
|
|
|
volumes:
|
|
dozzle_data:
|
|
dozzle_agent_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/kokoro/compose.yaml
|
|
name: kokoro-tts
|
|
services:
|
|
kokoro-tts:
|
|
container_name: kokoro-tts
|
|
build:
|
|
context: ./Kokoro-FastAPI
|
|
dockerfile: docker/gpu/Dockerfile
|
|
volumes:
|
|
- /tank/kokoro/models:/app/api/src/models
|
|
- /tank/kokoro/output:/app/output
|
|
ports:
|
|
- "8765:8880"
|
|
environment:
|
|
- PYTHONPATH=/app:/app/api
|
|
- USE_GPU=true
|
|
- PYTHONUNBUFFERED=1
|
|
- DOWNLOAD_MODEL=false
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities: [gpu]
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Kokoro TTS
|
|
- homepage.icon=mdi-waveform
|
|
- homepage.description=Kokoro FastAPI TTS (OpenAI-compatible)
|
|
- homepage.href=http://10.250.50.54:8765
|
|
restart: unless-stopped
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
|
|
>>> /opt/docker/compose/llama-swap/compose.yaml
|
|
# llama-swap — GGUF model server with on-demand model swapping.
|
|
#
|
|
# Proxies OpenAI-compatible API requests to llama.cpp server instances
|
|
# and swaps which model is loaded into VRAM per request. Runs on
|
|
# ana-ml2 using both GPUs dynamically (no explicit device pinning —
|
|
# llama-swap picks per-model-definition).
|
|
#
|
|
# Model definitions live in /opt/docker/conf/llama-swap/config.yaml on
|
|
# the server. Canonical copy of that config is config.yaml in this
|
|
# workspace; deploy with scp + `docker compose restart` or the script
|
|
# at the bottom of README.md.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
llama-swap:
|
|
image: ghcr.io/mostlygeek/llama-swap:${LLAMA_SWAP_VERSION}
|
|
container_name: llama-swap
|
|
restart: unless-stopped
|
|
stdin_open: true
|
|
tty: true
|
|
runtime: nvidia
|
|
ports:
|
|
- "${LLAMA_SWAP_PORT}:8080"
|
|
volumes:
|
|
- /opt/docker/conf/llama-swap/config.yaml:/app/config.yaml
|
|
- ${MODELS_DIR}:/models
|
|
- ${HF_CACHE_DIR}:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
# Pin to GPU 0 — the reserved card for on-demand large-model hot-loads.
|
|
# The always-on vLLM services (granite + embed/rerank/reward) own GPU 1;
|
|
# keeping llama-swap off GPU 1 stops a hot-loaded model from contending
|
|
# with them. llama.cpp then sees only GPU 0 (cuda:0), so --n-gpu-layers
|
|
# 999 loads there with no per-model device targeting needed.
|
|
- NVIDIA_VISIBLE_DEVICES=${LLAMA_SWAP_GPU:-0}
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/ >/dev/null || exit 1"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 30s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=llama-swap
|
|
- homepage.icon=mdi-swap-horizontal
|
|
- homepage.description=GGUF model swapper (llama.cpp; ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${LLAMA_SWAP_PORT}
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/parakeet/compose.yaml
|
|
services:
|
|
parakeet-stt:
|
|
image: parakeet-stt
|
|
ports:
|
|
- 8300:8000
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
restart: unless-stopped
|
|
volumes:
|
|
- parakeet_cache:/root/.cache
|
|
networks:
|
|
- tnet
|
|
env_file:
|
|
- .env
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Parakeet
|
|
- homepage.icon=mdi-talk
|
|
- homepage.description=Parakeet STT (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:8300
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
volumes:
|
|
parakeet_cache: null
|
|
|
|
>>> /opt/docker/compose/qwen35-vl/compose.yaml
|
|
# qwen35-vl — Qwen3.5-9B vision-language model (FP8) on ana-ml2.
|
|
#
|
|
# Co-located on GPU 1 with the granite summarizer + embed/rerank/reward trio
|
|
# (GPU 0 is deliberately kept free for hot-reloading large models). Serves on
|
|
# :8007, fronted by the LiteLLM gateway as `qwen3.5-9b-fp8`.
|
|
#
|
|
# WHY A PINNED NIGHTLY DIGEST (not :latest): vLLM :latest (v0.19.1) quantizes
|
|
# the Qwen3.5-VL *vision tower* under --quantization fp8, producing garbage
|
|
# vision output (the language model is unaffected — it answers text fine but
|
|
# "sees" noise). The nightly correctly excludes the vision tower, so vision
|
|
# works while the LM still gets the FP8 throughput/VRAM win. We pin the exact
|
|
# nightly digest for reproducibility — a moving :nightly tag would silently
|
|
# change the engine. WATCH: once the vision-FP8 exclusion lands in a stable
|
|
# release, re-pin to :latest and drop this note.
|
|
#
|
|
# WHY util 0.40 (not the trio's tiny values): the model needs ~34 GB just to
|
|
# start at 32k context (FP8 weights + BF16 vision tower + graph capture + 32k
|
|
# profiling). On shared GPU 1 (prod uses ~46 GB, ~48 GB free) this vLLM build
|
|
# requires free >= util*total, capping util at ~0.51 here; 0.40 (~38 GB) sits
|
|
# above the ~34 GB floor with ~10 GB card headroom.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
name: qwen35-vl
|
|
|
|
services:
|
|
vllm-qwen35:
|
|
image: ${QWEN_IMAGE}
|
|
container_name: ${QWEN_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QWEN_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${QWEN_MODEL}
|
|
- --served-model-name
|
|
- ${QWEN_SERVED_NAME}
|
|
- --quantization
|
|
- fp8
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${QWEN_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${QWEN_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
# Prefix caching pinned ON (the nightly defaults it OFF). Free win for the
|
|
# text-chat path; marginal for vision (each image is a distinct prefix).
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QWEN_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Qwen3.5-9B VL (FP8)
|
|
- homepage.icon=mdi-image-search
|
|
- homepage.description=Qwen3.5-9B vision-language (FP8) via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/vibevoice/compose.yaml
|
|
services:
|
|
vibevoice:
|
|
container_name: vibevoice
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
ports:
|
|
- 8745:8745
|
|
volumes:
|
|
- /tank/vibevoice/hf:/root/.cache/huggingface
|
|
- /tank/vibevoice/voices:/app/voices
|
|
- /tank/vibevoice/state:/var/lib/eworker
|
|
environment:
|
|
- ENABLE_1_5B=true
|
|
- ENABLE_LARGE=true
|
|
- AUTH_REQUIRED=true
|
|
- CORS_ENABLED=true
|
|
- ALLOWED_ORIGINS=*
|
|
image: eworkerinc/vibevoice:latest
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=VibeVoice
|
|
- homepage.icon=mdi-chat
|
|
- homepage.description=EWorkerStudio VibeVoice
|
|
- homepage.href=http://10.250.50.54:8745
|
|
restart: unless-stopped
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/vllm/compose.yaml
|
|
# vLLM — Qwen3 Embedding + Reranker + Skywork Reward-V2 classifier.
|
|
#
|
|
# Originally created to replace the unmaintained Infinity stack (embed +
|
|
# rerank); generalized 2026-05-13 to host any vLLM-served model on ana-ml2,
|
|
# starting with the Skywork-Reward-V2-Llama-3.1-8B reward classifier
|
|
# (AWQ-quantized locally, served from /tank/aimodels/llm/).
|
|
#
|
|
# vLLM runs one model per process, so this stack brings up three containers
|
|
# sharing a single GPU:
|
|
#
|
|
# vllm-embed — Qwen3-Embedding served as an OpenAI /v1/embeddings server
|
|
# vllm-rerank — Qwen3-Reranker served as a /rerank + /score server
|
|
# vllm-reward — Skywork-Reward-V2-Llama-3.1-8B-AWQ served as a /classify scorer
|
|
#
|
|
# The reranker is a causal-LM checkpoint; --hf-overrides re-maps it to
|
|
# Qwen3ForSequenceClassification so vLLM's reranking endpoints work and the
|
|
# model only emits two class logits (no/yes) instead of the full 151k vocab.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
#
|
|
# Pre-download models to avoid first-run delay:
|
|
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
|
|
# --var hf_repo=Qwen/Qwen3-Embedding-0.6B
|
|
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
|
|
# --var hf_repo=Qwen/Qwen3-Reranker-0.6B
|
|
#
|
|
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
|
|
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on ana-ml2 and is
|
|
# bind-mounted into the reward service at /local-models. Not from HF Hub.
|
|
|
|
services:
|
|
vllm-embed:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-embed
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${EMBED_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${EMBED_MODEL}
|
|
- --served-model-name
|
|
- ${EMBED_MODEL}
|
|
- --runner
|
|
- pooling
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${EMBED_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${EMBED_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=vLLM Embed (Qwen3)
|
|
- homepage.icon=mdi-vector-arrange-below
|
|
- homepage.description=Qwen3 Embedding via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${EMBED_PORT}/docs
|
|
|
|
vllm-rerank:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-rerank
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${RERANK_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${RERANK_MODEL}
|
|
- --served-model-name
|
|
- ${RERANK_MODEL}
|
|
- --runner
|
|
- pooling
|
|
- --hf-overrides
|
|
- '{"architectures":["Qwen3ForSequenceClassification"],"classifier_from_token":["no","yes"],"is_original_qwen3_reranker":true}'
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${RERANK_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${RERANK_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=vLLM Rerank (Qwen3)
|
|
- homepage.icon=mdi-sort-variant
|
|
- homepage.description=Qwen3 Reranker via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${RERANK_PORT}/docs
|
|
|
|
vllm-reward:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-reward
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${REWARD_PORT}:8000"
|
|
volumes:
|
|
# AWQ output lives in the legacy llama-swap models tree, not the HF cache
|
|
# — bind-mount the LLM models dir read-only so the reward service can
|
|
# load it as a local-path HF-format model.
|
|
- /tank/aimodels/llm:/local-models:ro
|
|
environment:
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- /local-models/Skywork-Reward-V2-Llama-3.1-8B-AWQ
|
|
- --served-model-name
|
|
- Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
|
|
# vLLM 0.19.1 deprecated --task in favor of --runner. The model's
|
|
# config.json declares `LlamaForSequenceClassification` so the
|
|
# pooling runner uses it as a classifier (single-label reward score)
|
|
# without needing an explicit task flag.
|
|
- --runner
|
|
- pooling
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${REWARD_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${REWARD_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 240s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=vLLM Reward (Skywork)
|
|
- homepage.icon=mdi-scale-balance
|
|
- homepage.description=Skywork-Reward-V2 8B classifier via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${REWARD_PORT}/docs
|
|
|
|
# Phi-4-mini (FP8) — summarizer + "dreaming" agent. Supersedes the
|
|
# llama-swap granite-4-small pin. Generative chat model (OpenAI
|
|
# /v1/chat/completions), so NO --runner pooling. FP8 on RTX 6000 Ada
|
|
# (cc 8.9): near-lossless, ~1.2x, ~6 GB.
|
|
vllm-granite:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-granite
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${GRANITE_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
# Production summarizer (replaced phi4-mini 2026-06-05). Default = official
|
|
# IBM pre-quantized FP8 (compressed-tensors), loaded directly; FP8 is native
|
|
# on the RTX 6000 Ada (cc 8.9). Fallback to vLLM-native dynamic FP8 from
|
|
# BF16: GRANITE_MODEL=ibm-granite/granite-4.1-8b + GRANITE_QUANT=fp8.
|
|
- ${GRANITE_MODEL}
|
|
- --served-model-name
|
|
- ${GRANITE_SERVED_NAME}
|
|
- --quantization
|
|
- ${GRANITE_QUANT}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${GRANITE_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${GRANITE_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
# CUDA graphs ENABLED (no --enforce-eager) for decode throughput. Made
|
|
# room 2026-06-05 by right-sizing the embed/rerank/reward trio's KV pools
|
|
# (they were over-provisioned at 5.9x/2.0x/3.9x concurrency); GPU 1 now has
|
|
# ~17 GB free after granite, so graph-capture buffers fit. If the trio
|
|
# ever grows back, granite may need --enforce-eager again on this card.
|
|
# FP8 KV cache — halves KV memory; near-lossless on Ada (cc 8.9).
|
|
- --kv-cache-dtype
|
|
- ${GRANITE_KV_CACHE_DTYPE}
|
|
# Prefix caching pinned EXPLICIT (vLLM v1 defaults it on, but pin so a
|
|
# version flip can't silently disable it). Benched 2026-06-13: ~6.5x faster
|
|
# TTFT (45ms vs 292ms) on a shared ~4.5k-token summarizer template; soft/
|
|
# evictable KV, neutral when prefixes don't repeat — pure win for granite.
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GRANITE_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=vLLM Granite 4.1 8B (summarizer)
|
|
- homepage.icon=mdi-text-box-outline
|
|
- homepage.description=Granite 4.1 8B FP8 via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${GRANITE_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
===== CONFIG LAYOUT (/opt/docker/conf/ — top 200 entries) =====
|
|
|
|
/opt/docker/conf
|
|
/opt/docker/conf/llama-swap
|
|
/opt/docker/conf/llama-swap/config.yaml
|
|
/opt/docker/conf/vllm
|
|
|
|
===== LISTENING PORTS =====
|
|
|
|
0.0.0.0:111
|
|
0.0.0.0:22
|
|
0.0.0.0:5001
|
|
0.0.0.0:7007
|
|
0.0.0.0:8001
|
|
0.0.0.0:8002
|
|
0.0.0.0:8003
|
|
0.0.0.0:8004
|
|
0.0.0.0:8007
|
|
0.0.0.0:9292
|
|
[::]:111
|
|
[::]:22
|
|
*:2375
|
|
[::]:5001
|
|
[::]:8001
|
|
[::]:8002
|
|
[::]:8003
|
|
[::]:8004
|
|
[::]:8007
|
|
[::]:9292
|
|
|
|
===== MODEL / HUGGINGFACE CACHES =====
|
|
|
|
/tank/aimodels/huggingface (342G)
|
|
hub entries:
|
|
CACHEDIR.TAG
|
|
datasets--HuggingFaceH4--ultrachat_200k
|
|
datasets--mlabonne--harmful_behaviors
|
|
datasets--mlabonne--harmless_alpaca
|
|
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
|
|
datasets--wikitext
|
|
models--AxionML--Qwen3.5-9B-NVFP4
|
|
models--bartowski--Meta-Llama-3.1-8B-Instruct-GGUF
|
|
models--bartowski--NousResearch_Hermes-4-14B-GGUF
|
|
models--bartowski--TheDrummer_GLM-Steam-106B-A12B-v1-GGUF
|
|
models--bartowski--TheDrummer_Skyfall-31B-v4-GGUF
|
|
models--BeaverAI--Artemis-31B-v1i-GGUF
|
|
models--BeaverAI--Skyfall-R1-31B-v4a-GGUF
|
|
models--drawais--Granite-4.1-30B-NVFP4
|
|
models--ibm-granite--granite-4.0-h-small-GGUF
|
|
models--ibm-granite--granite-4.0-h-tiny-GGUF
|
|
models--ibm-granite--granite-4.0-micro-GGUF
|
|
models--ibm-granite--granite-4.1-8b
|
|
models--ibm-granite--granite-4.1-8b-fp8
|
|
models--llmfan46--Qwen3.6-35B-A3B-uncensored-heretic-GGUF
|
|
models--microsoft--Phi-4-mini-instruct
|
|
models--mradermacher--Daredevil-8B-abliterated-dpomix-GGUF
|
|
models--mradermacher--Qwen3-30B-A3B-abliterated-erotic-i1-GGUF
|
|
models--mradermacher--Qwen3.6-35B-A3B-abliterated-i1-GGUF
|
|
models--mradermacher--Selene-1-Mini-Llama-3.1-8B-i1-GGUF
|
|
models--murilonwt--granite-4.1-8b-NVFP4
|
|
models--newsletter--VibeVoice-Large-pt
|
|
models--Qwen--Qwen2.5-0.5B-Instruct
|
|
models--Qwen--Qwen3.5-9B
|
|
models--Qwen--Qwen3.6-35B-A3B
|
|
|
|
/tank/aimodels/llm (794G)
|
|
|
|
/home/lkraven/.cache/huggingface (15G)
|
|
hub entries:
|
|
CACHEDIR.TAG
|
|
datasets--Salesforce--wikitext
|
|
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
|
|
datasets--wikitext
|
|
models--Astralyra--bge-reranker-large-Q8_0-GGUF
|
|
models--ggml-org--embeddinggemma-300M-GGUF
|
|
models--ggml-org--Qwen3-Reranker-0.6B-Q8_0-GGUF
|
|
models--jinaai--jina-reranker-v3-GGUF
|
|
models--klnstpr--bge-reranker-v2-m3-Q8_0-GGUF
|
|
models--minhtd14--jina-reranker-v2-base-multilingual-Q8_0-GGUF
|
|
models--Mungert--Qwen3-Reranker-0.6B-GGUF
|
|
models--Qwen--Qwen3-Embedding-0.6B-GGUF
|
|
models--Qwen--Qwen3-Reranker-0.6B
|
|
models--Skywork--Skywork-Reward-V2-Llama-3.1-8B
|
|
models--unsloth--gemma-4-26B-A4B-it
|
|
models--unsloth--gemma-4-26B-A4B-it-GGUF
|
|
models--unsloth--gemma-4-31B-it-GGUF
|
|
|
|
|
|
===== DOCKER-ADJACENT SYSTEMD SERVICES =====
|
|
|
|
containerd.service running
|
|
docker.service running
|
|
nvidia-persistenced.service running
|
|
|
|
===== DONE =====
|
|
|
|
Paste the above back into the chat, or pass a path as argv[1] to save.
|