966324c5f1
The README's running-stacks table had drifted well behind reality (still listed llama-swap + only the embed/rerank/reward trio). Regenerated system-details.txt and rewrote the stacks + GPU-allocation sections from a live docker ps + nvidia-smi (2026-07-22): - GPU 0 (hot): vllm-aeon-gen (qwen3.6-35b-a3b-heretic, NVFP4) + vllm-charrp- reasoning-nvfp4 (char-rp-reasoning, NVFP4), ~42+45 GB. - GPU 1 (on-demand): granite-4.1-8b, selene-1-mini-8b, Skywork reward, Qwen3 embed/rerank, and the Magidonia-24B char-RP GGUF (llama-charrp), ~91 GB. - Recorded the dormant on-disk stacks and llama-swap's retirement.
2494 lines
100 KiB
Plaintext
2494 lines
100 KiB
Plaintext
|
|
===== HOST =====
|
|
|
|
Hostname: ana-ml2
|
|
Date: 2026-07-22T15:24:06-07:00
|
|
Uptime: up 5 weeks, 5 days, 1 hour, 55 minutes
|
|
OS: Debian GNU/Linux 13 (trixie)
|
|
Kernel: 6.12.74+deb13+1-amd64
|
|
Arch: x86_64
|
|
|
|
===== HARDWARE =====
|
|
|
|
CPU cores: 96
|
|
CPU model: AMD EPYC 9254 24-Core Processor
|
|
MemTotal: 566.6 GB
|
|
MemAvailable: 229.0 GB
|
|
|
|
===== GPUS =====
|
|
|
|
index, name, memory.total [MiB], memory.free [MiB], driver_version
|
|
0, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 9494 MiB, 580.65.06
|
|
1, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 6238 MiB, 580.65.06
|
|
|
|
===== FILESYSTEMS (df) =====
|
|
|
|
Filesystem Size Used Avail Use% Mounted on
|
|
zroot/ROOT/debian 384G 271G 113G 71% /
|
|
efivarfs 128K 67K 57K 55% /sys/firmware/efi/efivars
|
|
/dev/sdb1 511M 92M 420M 18% /boot/efi
|
|
tank 8.6T 3.9T 4.8T 46% /tank
|
|
zroot/home 185G 72G 113G 40% /home
|
|
|
|
===== PERSISTENT MOUNTS (/etc/fstab, non-comment) =====
|
|
|
|
UUID="3D9B-8E0C" /boot/efi vfat defaults 0 0
|
|
|
|
===== TARGETED DATA PATHS =====
|
|
|
|
/tank (total: 3.9T)
|
|
total 20
|
|
drwxrwxrwx 7 root root 7 2026-06-14 16:11 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
drwxrwxr-x 31 llmuser llm 38 2026-07-14 09:27 aimodels
|
|
drwxrwxr-x 3 llmuser llm 3 2025-09-08 13:21 comfy
|
|
drwxrwxr-x 4 llmuser llmuser 4 2026-04-11 23:17 kokoro
|
|
drwxr-xr-x 5 r18clip r18clip 6 2026-06-14 16:20 r18-clip-caption
|
|
drwxrwxr-x 5 lkraven lkraven 6 2025-09-10 18:35 vibevoice
|
|
|
|
/opt (total: 16G)
|
|
total 79
|
|
drwxrwxrwx 13 root root 13 2026-04-17 23:00 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
drwx--x--x 4 root root 4 2025-09-02 12:53 containerd
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 docker
|
|
drwxrwxr-x 8 llmuser llm 17 2026-04-17 17:50 heretic
|
|
drwxrwxr-x 15 llmuser llmuser 36 2026-04-11 23:24 Kokoro-FastAPI
|
|
drwxrwxr-x 20 llmuser llmuser 39 2025-10-09 21:42 LibreChat
|
|
drwxrwxr-x 26 llmuser llmuser 56 2025-09-04 21:01 llama.cpp
|
|
drwxrwxr-x 13 llmuser llmuser 25 2025-09-02 12:50 llama-swap
|
|
drwxrwxr-x 3 llmuser llmuser 11 2026-04-18 15:48 llmcompressor
|
|
drwxr-xr-x 4 root root 4 2025-08-30 22:54 nvidia
|
|
drwxrwxr-x 6 llmuser llmuser 15 2025-09-24 09:53 parakeet-tdt-0.6b-v2-fastapi
|
|
drwxrwxr-x 2 llmuser llmuser 6 2025-09-05 10:03 uv
|
|
|
|
/opt/docker (total: 110M)
|
|
total 18
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 .
|
|
drwxrwxrwx 13 root root 13 2026-04-17 23:00 ..
|
|
drwxrwxr-x 24 lkraven lkraven 24 2026-07-14 14:50 compose
|
|
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 conf
|
|
|
|
/opt/docker/compose (total: 110M)
|
|
total 68
|
|
drwxrwxr-x 24 lkraven lkraven 24 2026-07-14 14:50 .
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 18:47 beszel-agent-ana
|
|
drwxrwxr-x 2 lkraven lkraven 8 2026-07-16 08:59 char-rp-gguf
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-05 20:36 comfyui
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:47 dockge
|
|
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:44 dozzle-agent-ana
|
|
drwxrwxr-x 3 lkraven lkraven 6 2026-07-16 09:13 heretic2-charrp-reasoning
|
|
drwxrwxr-x 3 lkraven lkraven 4 2026-04-11 23:17 kokoro
|
|
drwxr-xr-x 2 lkraven lkraven 6 2026-06-12 15:17 llama-swap
|
|
drwxr-xr-x 2 lkraven lkraven 6 2026-06-18 23:40 mistral-medium-3.5
|
|
drwxr-xr-x 2 lkraven lkraven 5 2026-06-15 18:40 mistral-small-4
|
|
drwxr-xr-x 2 root root 4 2026-06-17 22:30 mistral-small-4-heretic
|
|
drwxr-xr-x 2 root root 4 2026-07-08 01:10 ms32-24b-angel
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-24 12:32 parakeet
|
|
drwxr-xr-x 2 lkraven lkraven 5 2026-06-19 01:52 qwen3.5-122b
|
|
drwxrwxr-x 2 lkraven lkraven 4 2026-06-13 08:34 qwen35-vl
|
|
drwxrwxr-x 2 lkraven lkraven 8 2026-07-16 09:17 qwen36-27b-aeon
|
|
drwxr-xr-x 2 lkraven lkraven 5 2026-06-15 20:30 qwen36-vl
|
|
drwxr-xr-x 2 lkraven lkraven 5 2026-07-14 16:43 qwen-image-bench
|
|
drwxr-xr-x 2 lkraven lkraven 6 2026-07-14 19:56 qwopus3.5-122b
|
|
drwxr-xr-x 2 lkraven lkraven 5 2026-07-14 16:43 selene
|
|
drwxr-xr-x 2 lkraven lkraven 4 2025-09-10 18:03 vibevoice
|
|
drwxr-xr-x 2 lkraven lkraven 13 2026-07-16 10:45 vllm
|
|
|
|
/opt/docker/conf (total: 14K)
|
|
total 2
|
|
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 .
|
|
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
|
|
drwxr-xr-x 2 lkraven lkraven 3 2026-06-05 00:22 llama-swap
|
|
drwxr-xr-x 2 lkraven lkraven 2 2026-06-04 00:37 vllm
|
|
|
|
/var/lib/docker (total: 8.5K)
|
|
|
|
/srv (total: 1.0K)
|
|
total 10
|
|
drwxr-xr-x 2 root root 3 2026-06-14 16:11 .
|
|
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
|
|
lrwxrwxrwx 1 root root 22 2026-06-14 16:11 r18-clip-caption -> /tank/r18-clip-caption
|
|
|
|
|
|
===== DOCKER =====
|
|
|
|
Server: 29.3.1 Client: 29.3.1
|
|
|
|
----- docker info -----
|
|
Containers: 24 (running 11, paused 0, stopped 13)
|
|
Images: 53
|
|
Runtimes: map[io.containerd.runc.v2:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} nvidia:{{nvidia-container-runtime [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} runc:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]}]
|
|
Default runtime: runc
|
|
Storage driver: overlay2
|
|
Root dir: /var/lib/docker
|
|
Server version: 29.3.1
|
|
|
|
----- running containers -----
|
|
NAMES IMAGE STATUS PORTS
|
|
vllm-granite vllm/vllm-openai:latest Up 6 days (healthy) 0.0.0.0:8004->8000/tcp, [::]:8004->8000/tcp
|
|
vllm-aeon-gen vllm/vllm-openai:latest Up 6 days (healthy) 0.0.0.0:8015->8000/tcp, [::]:8015->8000/tcp
|
|
vllm-charrp-reasoning-nvfp4 vllm/vllm-openai:v0.24.0 Up 6 days (healthy) 0.0.0.0:8018->8000/tcp, [::]:8018->8000/tcp
|
|
llama-charrp ghcr.io/mostlygeek/llama-swap:cuda Up 6 days (healthy) 0.0.0.0:8016->8080/tcp, [::]:8016->8080/tcp
|
|
vllm-reward vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8003->8000/tcp, [::]:8003->8000/tcp
|
|
vllm-embed vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8001->8000/tcp, [::]:8001->8000/tcp
|
|
vllm-rerank vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8002->8000/tcp, [::]:8002->8000/tcp
|
|
vllm-selene vllm/vllm-openai Up 7 days (healthy) 0.0.0.0:8011->8000/tcp, [::]:8011->8000/tcp
|
|
dockge louislam/dockge:latest Up 5 weeks (healthy) 0.0.0.0:5001->5001/tcp, [::]:5001->5001/tcp
|
|
dozzle-agent amir20/dozzle:latest Up 5 weeks 0.0.0.0:7007->7007/tcp, 8080/tcp
|
|
beszel-agent henrygd/beszel-agent:latest Up 5 weeks (healthy)
|
|
|
|
----- all containers -----
|
|
NAMES IMAGE STATUS
|
|
vllm-granite vllm/vllm-openai:latest Up 6 days (healthy)
|
|
vllm-aeon-gen vllm/vllm-openai:latest Up 6 days (healthy)
|
|
vllm-charrp-reasoning-nvfp4 vllm/vllm-openai:v0.24.0 Up 6 days (healthy)
|
|
llama-charrp ghcr.io/mostlygeek/llama-swap:cuda Up 6 days (healthy)
|
|
vllm-qwopus35-122b vllm/vllm-openai:latest Created
|
|
vllm-aeon-rp vllm/vllm-openai:latest Created
|
|
llama-charrp-reasoning llamacpp-charrp:custom-latest Created
|
|
vllm-qwen-image-bench vllm/vllm-openai:latest Exited (0) 6 days ago
|
|
vllm-reward vllm/vllm-openai:latest Up 7 days (healthy)
|
|
vllm-embed vllm/vllm-openai:latest Up 7 days (healthy)
|
|
vllm-rerank vllm/vllm-openai:latest Up 7 days (healthy)
|
|
vllm-selene vllm/vllm-openai Up 7 days (healthy)
|
|
vllm-heretic2-modelopt-quant vllm/vllm-openai:v0.24.0 Exited (0) 8 days ago
|
|
vllm-heretic2-cg-quant vllm/vllm-openai:v0.24.0 Exited (0) 8 days ago
|
|
rp-dl6 vllm/vllm-openai:latest Exited (0) 2 weeks ago
|
|
rp-dl5 vllm/vllm-openai:latest Exited (0) 2 weeks ago
|
|
aeon-t1-sft aeon-trainer:latest Exited (0) 2 weeks ago
|
|
vllm-deckard-40b vllm/vllm-openai:v0.23.0 Exited (0) 3 weeks ago
|
|
selene-dl vllm/vllm-openai:v0.22.0 Exited (0) 5 weeks ago
|
|
mistral-dl f37691f675bb Exited (0) 5 weeks ago
|
|
r18-staging-dl df7be4c4d818 Exited (0) 5 weeks ago
|
|
dockge louislam/dockge:latest Up 5 weeks (healthy)
|
|
dozzle-agent amir20/dozzle:latest Up 5 weeks
|
|
beszel-agent henrygd/beszel-agent:latest Up 5 weeks (healthy)
|
|
|
|
----- networks -----
|
|
NAME DRIVER SCOPE
|
|
bridge bridge local
|
|
host host local
|
|
kokoro-tts-gpu_default bridge local
|
|
librechat_default bridge local
|
|
llama-swap_default bridge local
|
|
none null local
|
|
traefik-net bridge local
|
|
|
|
----- networks (external, non-default — worth knowing for compose external: true) -----
|
|
kokoro-tts-gpu_default
|
|
librechat_default
|
|
llama-swap_default
|
|
traefik-net
|
|
|
|
----- named volumes -----
|
|
VOLUME NAME DRIVER
|
|
beszel-agent-ana_beszel_agent_data local
|
|
dockge_dockge_data local
|
|
dozzle-agent-ana_dozzle_agent_data local
|
|
parakeet_parakeet_cache local
|
|
searxng_searxng-data local
|
|
|
|
----- compose projects currently running -----
|
|
beszel-agent-ana
|
|
char-rp-gguf
|
|
dockge
|
|
dozzle-agent-ana
|
|
heretic2-charrp-reasoning
|
|
qwen36-27b-aeon
|
|
selene
|
|
vllm
|
|
|
|
===== COMPOSE FILES (/opt/docker/compose/) =====
|
|
|
|
|
|
>>> /opt/docker/compose/beszel-agent-ana/compose.yaml
|
|
# Beszel — lightweight server/container monitoring.
|
|
#
|
|
# Hub: single web UI with the SQLite store. Agents: per-host metric collectors
|
|
# that the hub pulls from over SSH.
|
|
#
|
|
# Multi-host layout via compose profiles:
|
|
# COMPOSE_PROFILES=hub → hub only (ana-docker)
|
|
# COMPOSE_PROFILES=hub,agent → hub + local agent on the same host
|
|
# COMPOSE_PROFILES=agent → agent only (ana-ml2, nh3-docker,
|
|
# esh-docker-vm, vm-esh-nas)
|
|
#
|
|
# The agent uses network_mode: host so it sees real host CPU/mem/net/disk
|
|
# counters rather than container-scoped ones — that's why it can't share
|
|
# the tnet network with the hub.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
beszel:
|
|
image: henrygd/beszel:${BESZEL_VERSION}
|
|
container_name: beszel
|
|
profiles: [hub]
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${BESZEL_PORT}:8090"
|
|
volumes:
|
|
- beszel_data:/beszel_data
|
|
healthcheck:
|
|
# Hub image is distroless — no wget/curl. Use the bundled `/beszel`
|
|
# binary's built-in health subcommand (https://beszel.dev/guide/healthchecks).
|
|
test: ["CMD", "/beszel", "health", "--url", "http://localhost:8090"]
|
|
interval: 120s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Monitoring
|
|
- homepage.name=Beszel
|
|
- homepage.icon=mdi-chart-line
|
|
- homepage.description=Server + container monitoring
|
|
- homepage.href=http://10.250.50.70:${BESZEL_PORT}
|
|
|
|
beszel-agent:
|
|
image: henrygd/beszel-agent:${BESZEL_VERSION}
|
|
container_name: beszel-agent
|
|
profiles: [agent]
|
|
restart: unless-stopped
|
|
network_mode: host
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- beszel_agent_data:/var/lib/beszel-agent
|
|
environment:
|
|
# Agent auth has two modes (v0.13+ supports both side-by-side):
|
|
# - KEY-mode: agent listens, hub connects inbound over SSH using KEY.
|
|
# Requires BESZEL_HUB_KEY in .env.
|
|
# - Token-mode: agent initiates an outbound connection to HUB_URL
|
|
# using TOKEN. Easier through NAT. Requires HUB_URL + BESZEL_TOKEN.
|
|
# Leave unused ones empty ("") in .env; both can be set simultaneously.
|
|
- PORT=${BESZEL_AGENT_PORT:-45876}
|
|
- KEY=${BESZEL_HUB_KEY:-}
|
|
- HUB_URL=${HUB_URL:-}
|
|
- TOKEN=${BESZEL_TOKEN:-}
|
|
- EXTRA_FILESYSTEMS=${BESZEL_EXTRA_FS:-}
|
|
healthcheck:
|
|
# Agent image ships the `/agent` binary with a `health` subcommand.
|
|
# Verifies the agent process is up — not that the hub can reach it.
|
|
test: ["CMD", "/agent", "health"]
|
|
interval: 120s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
|
|
volumes:
|
|
beszel_data:
|
|
beszel_agent_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/char-rp-gguf/compose.yaml
|
|
# char-rp-gguf — dedicated GGUF character-RP seat on ana-ml2 GPU 0, REPLACING the
|
|
# broken ms32-24b-angel NVFP4 serve (garbage output — bad self-quant W4A4).
|
|
#
|
|
# Two co-located llama.cpp (llama-server) instances on GPU 0, served alongside the
|
|
# 35B-A3B heretic `gen` (qwen36-27b-aeon stack, :8015):
|
|
#
|
|
# llama-charrp (:8016, gateway char-rp) — TheDrummer Magidonia-24B-v4.3 Q6_K.
|
|
# Magistral (Mistral) dark-romantasy RP tune. NON-thinking PROSE seat: elite
|
|
# literary prose, zero refusal, ~65 tok/s, precise POV/instruction adherence.
|
|
#
|
|
# llama-charrp-reasoning (:8018, gateway char-rp-reasoning) — ArliAI QwQ-32B-RpR-v4 Q5_K_M.
|
|
# QwQ reasoning RP tune whose reasoning DATA was generated with QwQ-ABLITERATED
|
|
# → it does NOT re-censor in the think phase (the exact failure mode that killed
|
|
# the Pantheon/DeepSeek-distilled reasoners: they reason themselves into refusals
|
|
# inside <think>). llama.cpp MANAGES QwQ reasoning natively: --reasoning on
|
|
# surfaces the trace in reasoning_content (clean prose in content, no <think>
|
|
# leak), --reasoning-budget caps the chain-of-thought. ~50 tok/s @ Q5_K_M.
|
|
#
|
|
# WHY GGUF/llama.cpp (not vLLM NVFP4): sidesteps BOTH traps that killed the Angel serve
|
|
# — the vLLM NVFP4 self-quant breakage AND the Mistral-tokenizer/vision crash. llama.cpp
|
|
# handles Mistral + QwQ tokenizers natively. NEVER Ollama (banned fleet-wide).
|
|
#
|
|
# WHY TWO models (not one): no single dense 24-32B is BOTH an elite non-thinking prose
|
|
# seat AND a clean managed-reasoning seat on llama.cpp. Magidonia's Magistral [THINK]
|
|
# discipline is loose (won't reliably close [/THINK] on substantive reasoning → prose
|
|
# bleeds into reasoning_content, content empties); Cydonia-R1's <think> is emergent, so
|
|
# llama.cpp can't manage/cap it → runaway CoT that never reaches prose. QwQ's template
|
|
# opens <think> natively → llama.cpp manages+caps it. So: best-of-breed per seat.
|
|
# ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning): point both services
|
|
# at Magidonia via CHARRP_REASONING_MODEL in .env and blank CHARRP_REASONING_EXTRA_*.
|
|
#
|
|
# ALTERNATE prose model: PaintedFantasy-v4.1-24B (also Magistral, more literary flair
|
|
# but looser POV adherence) — set CHARRP_MODEL in .env. All candidate GGUFs are
|
|
# pre-pulled to /tank/aimodels/llm/rp/.
|
|
#
|
|
# VRAM (GPU 0, co-resident with gen ~38G): Magidonia Q6 ~19G + RpR-v4 Q5 ~23G + KV/
|
|
# compute ~6-8G = ~85-88G / 97G (~9-12G margin). Keep ctx modest; drop CHARRP_*_CTX
|
|
# to 8192 in .env if warmup bites. depends_on sequences char-rp first.
|
|
#
|
|
# API auth: blank (LAN-internal on the GPU host; matches API_KEY= in the AEON stack /
|
|
# gateway VLLM_API_KEY). llama-server ignores the gateway's api_key when none is set.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
name: char-rp-gguf
|
|
|
|
services:
|
|
# ── PROSE seat — non-thinking. gateway char-rp. ──
|
|
llama-charrp:
|
|
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
|
container_name: ${CHARRP_CONTAINER:-llama-charrp}
|
|
restart: unless-stopped
|
|
runtime: nvidia
|
|
ports:
|
|
- "${CHARRP_PORT:-8016}:8080"
|
|
volumes:
|
|
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
|
|
environment:
|
|
# Pin to GPU 0 (the on-demand large-model card; the always-on vLLM trio owns GPU 1).
|
|
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
|
|
entrypoint: ["/app/llama-server"]
|
|
command:
|
|
- --model
|
|
- /models/${CHARRP_MODEL:-rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8080"
|
|
- --n-gpu-layers
|
|
- "999"
|
|
- --ctx-size
|
|
- "${CHARRP_CTX:-98304}"
|
|
- --flash-attn
|
|
- on
|
|
# q8_0 KV cache ~halves KV VRAM (8-bit, near-lossless) → ~2x the context per GB.
|
|
# Mistral/Magistral handles q8 KV cleanly. Set f16 in .env to disable.
|
|
- --cache-type-k
|
|
- ${CHARRP_KV_TYPE:-q8_0}
|
|
- --cache-type-v
|
|
- ${CHARRP_KV_TYPE:-q8_0}
|
|
- --jinja
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 240s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Inference
|
|
- homepage.name=char-rp (Magidonia-24B GGUF)
|
|
- homepage.icon=mdi-drama-masks
|
|
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
|
|
|
|
# ── REASONING seat — NEO-CODE = Heretic2-Thinking (Qwen3.6-27B) managed thinking. gateway char-rp-reasoning. ──
|
|
llama-charrp-reasoning:
|
|
# ⚠️ CUSTOM llama.cpp build (master 6eddde0 + unmerged PR #25544). Needed for two reasons:
|
|
# (1) recent master parses Qwen3.6's native qwen3_coder tool-call format (stock b8840 predates it —
|
|
# the <tool_call><function=..><parameter=..> XML is Qwen3.5/3.6-native, NOT an OpenHands quirk);
|
|
# (2) PR #25544 multi-terminator reasoning-budget fix (Worldtree #355) — belt-and-suspenders now that
|
|
# NEO-CODE shows 0.0 runaway (R36 gate), but keep it. DO NOT revert to stock until #25544 merges.
|
|
# Build recipe + why + rollback: ./llamacpp-custom/README.md.
|
|
# Rollback: set LLAMA_REASONING_IMAGE=ghcr.io/mostlygeek/llama-swap:cuda in .env + recreate.
|
|
image: ${LLAMA_REASONING_IMAGE:-llamacpp-charrp:custom-latest}
|
|
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
|
|
restart: unless-stopped
|
|
runtime: nvidia
|
|
# Sequence AFTER the prose seat is healthy so the two GPU-0 allocations don't race.
|
|
depends_on:
|
|
llama-charrp:
|
|
condition: service_healthy
|
|
ports:
|
|
- "${CHARRP_REASONING_PORT:-8018}:8080"
|
|
volumes:
|
|
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
|
|
environment:
|
|
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
|
|
entrypoint: ["/app/llama-server"]
|
|
command:
|
|
- --model
|
|
- /models/${CHARRP_REASONING_MODEL:-rp/Qwen3.6-27B-NEO-CODE-HERE-2T-OT-Q5_K_M.gguf}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8080"
|
|
- --n-gpu-layers
|
|
- "999"
|
|
- --ctx-size
|
|
- "${CHARRP_REASONING_CTX:-40960}"
|
|
- --flash-attn
|
|
- on
|
|
# NEO-CODE = Qwen3.6-27B GDN-hybrid (16 of 64 layers cache KV → KV cheap); native ctx 262144
|
|
# (256K). Full 256K @ q8_0 KV ≈ 8.6G, fits GPU0 w/ ~3.8G margin. q8_0 coherent; f16 in .env if gibberish.
|
|
- --cache-type-k
|
|
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
|
- --cache-type-v
|
|
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
|
- --jinja
|
|
# NEO-CODE's Qwen3.6 template natively opens <think> → llama.cpp manages the reasoning
|
|
# (trace to reasoning_content, content stays clean prose); --reasoning-budget caps the CoT.
|
|
# (R36 gate 2026-07-14: NEO-CODE composite 0.922 tool-calling + 0.0 runaway — beat Deckard
|
|
# 0.08/0.80 and gen-reasoning 0.856. Budget held at 400: latency-coupled to soong's client timeout.)
|
|
- --reasoning
|
|
- on
|
|
- --reasoning-format
|
|
- deepseek
|
|
- --reasoning-budget
|
|
- "${CHARRP_REASONING_BUDGET:-400}"
|
|
# Sampler defaults per the DavidAU/Qwen3.6 model card (thinking-mode, general tasks): temp 1.0,
|
|
# top_p 0.95, top_k 20, min_p 0.0, no rep-penalty, no DRY (DRY was a QwQ/Deckard looping band-aid
|
|
# NEO-CODE doesn't need). All tunable via .env. NOTE: 0.922 tool-gate was on the OLD Deckard
|
|
# samplers (effective temp~0.8 + DRY); re-validate tools + slop on these card samplers.
|
|
- --temp
|
|
- "${CHARRP_REASONING_TEMP:-1.0}"
|
|
- --top-p
|
|
- "${CHARRP_REASONING_TOP_P:-0.95}"
|
|
- --top-k
|
|
- "${CHARRP_REASONING_TOP_K:-20}"
|
|
- --min-p
|
|
- "${CHARRP_REASONING_MIN_P:-0.0}"
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Dormant
|
|
- homepage.name=char-rp-reasoning (QwQ-32B RpR-v4 GGUF)
|
|
- homepage.icon=mdi-brain
|
|
- homepage.description=Dark-romantasy RP reasoning seat, managed CoT (llama.cpp, ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${CHARRP_REASONING_PORT:-8018}
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/comfyui/compose.yaml
|
|
services:
|
|
comfyui:
|
|
runtime: nvidia
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
- compute
|
|
- utility
|
|
ports:
|
|
- 8188:8188
|
|
image: mmartial/comfyui-nvidia-docker:ubuntu24_cuda13.0-latest
|
|
networks:
|
|
- tnet
|
|
volumes:
|
|
- /tank/comfy/run:/comfy/mnt
|
|
- /tank/aimodels/img/comfy:/basedir
|
|
#user: 1001:1002
|
|
environment:
|
|
- WANTED_UID=1001
|
|
- WANTED_GID=1002
|
|
- BASE_DIRECTORY=/basedir
|
|
- SECURITY_LEVEL=weak
|
|
- NVIDIA_VISIBLE_DEVICES=all
|
|
- NVIDIA_DRIVER_CAPABILITIES=all
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=ComfyUI
|
|
- homepage.icon=mdi-panorama-variant-outline
|
|
- homepage.description=ComfyUI Image Gen (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:8188
|
|
restart: unless-stopped
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/dockge/compose.yaml
|
|
# Dockge — per-host Docker Compose UI (https://dockge.kuma.pet/).
|
|
#
|
|
# One instance runs on every Docker host so the compose dir is manageable
|
|
# from a browser. Each host sets DOCKGE_HOST_LABEL + DOCKGE_HOST_IP in its
|
|
# .env so the homepage card points at the right place.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
dockge:
|
|
image: louislam/dockge:${DOCKGE_VERSION:-latest}
|
|
container_name: dockge
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${DOCKGE_PORT:-5001}:5001"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock
|
|
- dockge_data:/app/data
|
|
- /opt/docker:/opt/docker
|
|
environment:
|
|
- DOCKGE_STACKS_DIR=/opt/docker/compose
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Service Networking
|
|
- homepage.name=Dockge (${DOCKGE_HOST_LABEL})
|
|
- homepage.icon=sh-dockge.png
|
|
- homepage.description=Compose UI on ${DOCKGE_HOST_LABEL}
|
|
- homepage.href=http://${DOCKGE_HOST_IP}:${DOCKGE_PORT:-5001}
|
|
|
|
volumes:
|
|
dockge_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/dozzle-agent-ana/compose.yaml
|
|
# Dozzle — container log viewer.
|
|
#
|
|
# Multi-host layout via compose profiles:
|
|
# COMPOSE_PROFILES=hub → runs the web UI (deploy on ana-docker)
|
|
# COMPOSE_PROFILES=agent → runs the remote agent (deploy on ana-ml2)
|
|
#
|
|
# Same compose.yaml on both servers; per-host `.env` picks the profile.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
dozzle:
|
|
image: amir20/dozzle:${DOZZLE_VERSION}
|
|
container_name: dozzle
|
|
profiles: [hub]
|
|
restart: unless-stopped
|
|
ports:
|
|
- "${DOZZLE_PORT}:8080"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- dozzle_data:/data
|
|
environment:
|
|
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
|
|
- DOZZLE_REMOTE_AGENT=${DOZZLE_REMOTE_AGENT:-}
|
|
- DOZZLE_AUTH_PROVIDER=${DOZZLE_AUTH_PROVIDER:-none}
|
|
- DOZZLE_USERNAME=${DOZZLE_USERNAME:-}
|
|
- DOZZLE_PASSWORD=${DOZZLE_PASSWORD:-}
|
|
healthcheck:
|
|
test: ["CMD", "/dozzle", "healthcheck"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 15s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=Monitoring
|
|
- homepage.name=Dozzle
|
|
- homepage.icon=mdi-text-box-search
|
|
- homepage.description=Container logs (ana-docker + ana-ml2)
|
|
- homepage.href=http://10.250.50.70:${DOZZLE_PORT}
|
|
|
|
dozzle-agent:
|
|
image: amir20/dozzle:${DOZZLE_VERSION}
|
|
container_name: dozzle-agent
|
|
profiles: [agent]
|
|
restart: unless-stopped
|
|
command: agent
|
|
ports:
|
|
- "${DOZZLE_AGENT_BIND:-0.0.0.0}:${DOZZLE_AGENT_PORT}:7007"
|
|
volumes:
|
|
- /var/run/docker.sock:/var/run/docker.sock:ro
|
|
- dozzle_agent_data:/data
|
|
environment:
|
|
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
|
|
networks:
|
|
- tnet
|
|
|
|
volumes:
|
|
dozzle_data:
|
|
dozzle_agent_data:
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/heretic2-charrp-reasoning/compose.yaml
|
|
# heretic2-charrp-reasoning — modelopt NVFP4 + native MTP fast char-rp-reasoning seat on
|
|
# ana-ml2 GPU0, replacing the GGUF NEO-CODE reasoning seat (llama-charrp-reasoning, now retired).
|
|
# Same Heretic2/NEO-CODE model; ~77 tok/s (~1.3x over GGUF) via qwen3_5_mtp spec-decode.
|
|
#
|
|
# ⚠️ REQUIRES the MTP workaround: vLLM 0.24.0 doesn't propagate modelopt exclude_modules to the
|
|
# spec-decode DRAFT model, so the BF16 mtp head gets quantized -> shape crash. The mounted
|
|
# sitecustomize.py (conf/mtp-workaround/) force-skips mtp.* in is_layer_skipped. Without it the
|
|
# engine dies at load. Full recipe: eshpfi docs/runbooks/heretic2-nvfp4-mtp-seat.md.
|
|
#
|
|
# Co-located on GPU0 with vllm-aeon-gen (gen) + llama-charrp (char-rp). VRAM: NVFP4 27B weights
|
|
# ~26GB + KV. util 0.30 fits the ~33GB free alongside gen+char-rp -> max-model-len capped at
|
|
# 32768 (the GGUF seat did 256K on lighter Q5 weights; NVFP4 is heavier, so context is reduced
|
|
# until VRAM is rebalanced). Tunables in .env.
|
|
|
|
name: heretic2-charrp-reasoning
|
|
|
|
services:
|
|
vllm-charrp-reasoning:
|
|
image: ${REASONING_IMAGE:-vllm/vllm-openai:v0.24.0}
|
|
container_name: ${REASONING_CONTAINER:-vllm-charrp-reasoning-nvfp4}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${REASONING_PORT:-8018}:8000"
|
|
volumes:
|
|
- /tank/aimodels:/tank/aimodels
|
|
# The MTP draft-model quant workaround (sitecustomize.py). PYTHONPATH loads it in the
|
|
# engine-core subprocess. See runbook landmine #4.
|
|
- ./conf/mtp-workaround:/mtp-workaround:ro
|
|
environment:
|
|
- PYTHONPATH=/mtp-workaround
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${REASONING_MODEL:-/tank/aimodels/heretic2-nvfp4-work/heretic2-modelopt-nvfp4-mtp}
|
|
- --quantization
|
|
- modelopt
|
|
- --speculative-config
|
|
- '{"method": "qwen3_5_mtp", "num_speculative_tokens": ${SPEC_TOKENS:-3}}'
|
|
- --language-model-only
|
|
- --mamba-cache-dtype
|
|
- float32
|
|
- --reasoning-parser
|
|
- qwen3
|
|
- --tool-call-parser
|
|
- qwen3_coder
|
|
- --enable-auto-tool-choice
|
|
- --served-model-name
|
|
- char-rp-reasoning
|
|
- --max-model-len
|
|
- "${REASONING_MAX_MODEL_LEN:-32768}"
|
|
- --max-num-seqs
|
|
- "${REASONING_MAX_NUM_SEQS:-4}"
|
|
- --gpu-memory-utilization
|
|
- "${REASONING_GPU_MEM_UTIL:-0.30}"
|
|
- --kv-cache-dtype
|
|
- fp8
|
|
- --trust-remote-code
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${REASONING_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 600s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Inference
|
|
- homepage.name=char-rp-reasoning (Heretic2 NVFP4+MTP)
|
|
- homepage.icon=mdi-rocket-launch
|
|
- homepage.description=NEO-CODE Heretic2 NVFP4 + native MTP, ~77 tok/s (ana-ml2 GPU0)
|
|
- homepage.href=http://10.250.50.54:${REASONING_PORT:-8018}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/kokoro/compose.yaml
|
|
name: kokoro-tts
|
|
services:
|
|
kokoro-tts:
|
|
container_name: kokoro-tts
|
|
build:
|
|
context: ./Kokoro-FastAPI
|
|
dockerfile: docker/gpu/Dockerfile
|
|
volumes:
|
|
- /tank/kokoro/models:/app/api/src/models
|
|
- /tank/kokoro/output:/app/output
|
|
ports:
|
|
- "8765:8880"
|
|
environment:
|
|
- PYTHONPATH=/app:/app/api
|
|
- USE_GPU=true
|
|
- PYTHONUNBUFFERED=1
|
|
- DOWNLOAD_MODEL=false
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities: [gpu]
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Kokoro TTS
|
|
- homepage.icon=mdi-waveform
|
|
- homepage.description=Kokoro FastAPI TTS (OpenAI-compatible)
|
|
- homepage.href=http://10.250.50.54:8765
|
|
restart: unless-stopped
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
|
|
>>> /opt/docker/compose/llama-swap/compose.yaml
|
|
# llama-swap — GGUF model server with on-demand model swapping.
|
|
#
|
|
# Proxies OpenAI-compatible API requests to llama.cpp server instances
|
|
# and swaps which model is loaded into VRAM per request. Runs on
|
|
# ana-ml2 using both GPUs dynamically (no explicit device pinning —
|
|
# llama-swap picks per-model-definition).
|
|
#
|
|
# Model definitions live in /opt/docker/conf/llama-swap/config.yaml on
|
|
# the server. Canonical copy of that config is config.yaml in this
|
|
# workspace; deploy with scp + `docker compose restart` or the script
|
|
# at the bottom of README.md.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
services:
|
|
llama-swap:
|
|
image: ghcr.io/mostlygeek/llama-swap:${LLAMA_SWAP_VERSION}
|
|
container_name: llama-swap
|
|
restart: unless-stopped
|
|
stdin_open: true
|
|
tty: true
|
|
runtime: nvidia
|
|
ports:
|
|
- "${LLAMA_SWAP_PORT}:8080"
|
|
volumes:
|
|
- /opt/docker/conf/llama-swap/config.yaml:/app/config.yaml
|
|
- ${MODELS_DIR}:/models
|
|
- ${HF_CACHE_DIR}:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
# Pin to GPU 0 — the reserved card for on-demand large-model hot-loads.
|
|
# The always-on vLLM services (granite + embed/rerank/reward) own GPU 1;
|
|
# keeping llama-swap off GPU 1 stops a hot-loaded model from contending
|
|
# with them. llama.cpp then sees only GPU 0 (cuda:0), so --n-gpu-layers
|
|
# 999 loads there with no per-model device targeting needed.
|
|
- NVIDIA_VISIBLE_DEVICES=${LLAMA_SWAP_GPU:-0}
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/ >/dev/null || exit 1"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 30s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=llama-swap
|
|
- homepage.icon=mdi-swap-horizontal
|
|
- homepage.description=GGUF model swapper (llama.cpp; ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${LLAMA_SWAP_PORT}
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/mistral-medium-3.5/compose.yaml
|
|
# mistral-medium-3.5 — RecViking/Mistral-Medium-3.5-128B-NVFP4 on ana-ml2 GPU 0
|
|
# as a TEMPORARY speed-check tenant, DISPLACING mistral-small-4 (operator 2026-06-19:
|
|
# "pull the NVFP4 model and serve it ... displace mistral-small-4 for now ... I want
|
|
# to check it").
|
|
#
|
|
# Model: Mistral Medium 3.5, 128B `Mistral3ForConditionalGeneration` (mistral3,
|
|
# multimodal), RecViking's NVFP4 (compressed-tensors / nvfp4-pack-quantized), HF format.
|
|
#
|
|
# SERVING (per RecViking's model card): vLLM NIGHTLY loads the HF-format NVFP4 weights
|
|
# DIRECTLY — no Mistral native-convert (unlike Small 4) — via the FlashInfer Cutlass
|
|
# NVFP4 kernel + TURBOQUANT 4-bit KV. RecViking used TP=4; the ~70 GB NVFP4 fits one
|
|
# 96 GB Blackwell card, so we run TP=1 on GPU 0. Context trimmed to 32K (speed check,
|
|
# not full 256K) so KV fits comfortably on one card.
|
|
#
|
|
# DISPLACEMENT: GPU 0 holds only one mistral-class model. Bring this up only after
|
|
# downing the live mistral-small-4-heretic stack. REVERT = `docker compose down` this,
|
|
# then `up -d` /opt/docker/compose/mistral-small-4-heretic (restores the Worldtree
|
|
# character backend). Served under its OWN name (mistral-medium-3.5), NOT
|
|
# mistral-small-4 — no stale-alias (the Worldtree character `mistral-small-4` route
|
|
# 404s while this is up; that's the "for now").
|
|
#
|
|
# EAGLE: RecViking's repo has no EAGLE head. The official native head
|
|
# (mistralai/Mistral-Medium-3.5-128B-EAGLE) is staged at /tank/aimodels/
|
|
# mistral-medium-3.5-eagle, but wiring spec-decode (native head + HF base, mistral3
|
|
# arch, nightly) is a follow-on — off by default. See README.
|
|
|
|
name: mistral-medium-3.5
|
|
|
|
services:
|
|
vllm-medium35:
|
|
image: ${MEDIUM35_IMAGE:-vllm/vllm-openai:nightly}
|
|
container_name: ${MEDIUM35_CONTAINER_NAME:-vllm-medium35}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${MEDIUM35_PORT:-8012}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
# RecViking NVFP4 checkpoint (HF format, read-only).
|
|
- /tank/aimodels/mistral-medium-3.5-nvfp4:/model:ro
|
|
# Official EAGLE draft head (native FP8, 2-layer) for speculative decoding.
|
|
- /tank/aimodels/mistral-medium-3.5-eagle:/eagle:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- /model
|
|
# HF-format load (NO --config-format/--load-format/--tokenizer-mode mistral —
|
|
# that's the Small 4 native path; nightly serves this NVFP4 from HF directly).
|
|
- --served-model-name
|
|
- mistral-medium-3.5
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --tensor-parallel-size
|
|
- "1"
|
|
- --gpu-memory-utilization
|
|
- ${MEDIUM35_GPU_MEM_UTIL:-0.93}
|
|
- --max-model-len
|
|
- ${MEDIUM35_MAX_MODEL_LEN:-32768}
|
|
# TURBOQUANT 4-bit KV (nightly) — RecViking's recommended KV dtype for this NVFP4.
|
|
- --kv-cache-dtype
|
|
- ${MEDIUM35_KV_CACHE_DTYPE:-turboquant_4bit_nc}
|
|
- --max-num-seqs
|
|
- ${MEDIUM35_MAX_NUM_SEQS:-16}
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
# EAGLE speculative decoding (method eagle, 3 spec tokens per the EAGLE card).
|
|
# Draft head mounted at /eagle. Remove these to revert to base-only.
|
|
# --enforce-eager: the EAGLE+NVFP4+turboquant path crashes in CUDA-graph replay
|
|
# on this nightly; disabling graphs is the workaround (costs some base speed).
|
|
- --enforce-eager
|
|
- --speculative-config
|
|
- '{"model": "/eagle", "num_speculative_tokens": 3, "method": "eagle"}'
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${MEDIUM35_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
# Nightly image + 70 GB NVFP4 load + CUDA/kernel warmup.
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Mistral Medium 3.5 (NVFP4 speed-check)
|
|
- homepage.icon=mdi-speedometer
|
|
- homepage.description=RecViking Mistral-Medium-3.5-128B NVFP4 on vLLM nightly (ana-ml2 GPU 0, displacing Small 4)
|
|
- homepage.href=http://10.250.50.54:${MEDIUM35_PORT:-8012}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/mistral-small-4/compose.yaml
|
|
# mistral-small-4 — Mistral-Small-4-119B-2603 (official NVFP4) on ana-ml2 GPU 0.
|
|
#
|
|
# Mistral Small 4 is a 119B-total / 6.5B-active MoE (128 experts, 4 active),
|
|
# 256K context, multimodal, Apache-2.0 (released 2026-03). This serves the
|
|
# OFFICIAL NVFP4 checkpoint (mistralai/Mistral-Small-4-119B-2603-NVFP4) — 74.4 GB
|
|
# of compressed-tensors (llm-compressor, a vLLM + Red Hat collaboration, day-0
|
|
# vLLM support). It is the GPU-0 tenant (the slot formerly reserved for a
|
|
# creative-writing pick — operator reassigned 2026-06-15; tune-for-creative-
|
|
# writing comes after base-characteristic probing).
|
|
#
|
|
# WHY NVFP4 (not FP8/bf16): on a SINGLE 96 GB card, NVFP4 (74.4 GB weights) is
|
|
# the only variant that fits at TP=1 — FP8 (~119 GB) and bf16 (~238 GB) need both
|
|
# GPUs. The card is Blackwell (sm_120) with FP4 tensor cores, so NVFP4 gets a real
|
|
# speedup, not just a VRAM save. NOTE: this is the COMPRESSED-TENSORS NVFP4 path
|
|
# (vendor-shipped, vLLM-tested) — distinct from the nvidia-ModelOpt NVFP4 MoE
|
|
# loader that broke on Qwen3.6 (#44081); different code path, day-0 supported.
|
|
#
|
|
# WHY TP=1 here: Mistral's official card uses --tensor-parallel-size 2 (their
|
|
# reference 80 GB cards can't fit 74.4 GB + context on one). The 96 GB Blackwell
|
|
# flips that to single-card: 74.4 GB weights + ~5 GB overhead leaves ~17 GB for
|
|
# KV. Mistral Small 4 uses MLA attention (TRITON_MLA) so KV is compressed/cheap —
|
|
# big context stays affordable even on a constrained KV pool. max-model-len is
|
|
# capped to 131072 on first bring-up (raise toward the native 256K once real KV
|
|
# headroom is measured).
|
|
#
|
|
# vLLM FLOOR: needs >= 0.20 (Mistral Small 4 day-0 support); validated on 0.23.0.
|
|
# Do NOT reuse the qwen36-vl 0.19.1 image — it predates this model.
|
|
#
|
|
# Serve flags mirror Mistral's official command (cited in README), adapted for
|
|
# single-card: TP 2->1, util 0.8->0.93, max-len 262144->131072, max-num-seqs
|
|
# 128->64. All tunables live in .env — edit that, not this file.
|
|
|
|
name: mistral-small-4
|
|
|
|
services:
|
|
vllm-mistral4:
|
|
image: ${MISTRAL_IMAGE}
|
|
container_name: ${MISTRAL_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${MISTRAL_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${MISTRAL_MODEL}
|
|
# Pre-quantized NVFP4 (compressed-tensors) — vLLM auto-detects the quant;
|
|
# no --quantization flag.
|
|
- --served-model-name
|
|
- mistral-small-4
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --tensor-parallel-size
|
|
- "1"
|
|
- --gpu-memory-utilization
|
|
- ${MISTRAL_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${MISTRAL_MAX_MODEL_LEN}
|
|
# MLA attention backend (DeepSeek-style latent KV → compressed, cheap KV).
|
|
- --attention-backend
|
|
- TRITON_MLA
|
|
# Mistral tool-calling + configurable reasoning (per the official card).
|
|
- --tool-call-parser
|
|
- mistral
|
|
- --enable-auto-tool-choice
|
|
- --reasoning-parser
|
|
- mistral
|
|
- --max-num-seqs
|
|
- ${MISTRAL_MAX_NUM_SEQS}
|
|
# VISION ENABLED. vLLM is pinned to v0.22.0 in .env — the last release BEFORE
|
|
# the Mistral multimodal regression (#44911, `MistralCommonImageProcessor has
|
|
# no attribute fetch_images`, landed ~0.22.1+; 0.23.0 is affected). v0.22.0
|
|
# still has Mistral-Small-4 arch + compressed-tensors NVFP4 support (the
|
|
# #44081 ModelOpt-NVFP4 bug on 0.22.0 is a DIFFERENT quant path, doesn't touch
|
|
# this compressed-tensors checkpoint). Gives a verified working vision tower
|
|
# as the abliteration/tuning baseline. (qwen36 stays on 0.23.0 — separate
|
|
# container; it NEEDS 0.23.0 for its ModelOpt NVFP4.)
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${MISTRAL_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 600s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Mistral Small 4 (NVFP4)
|
|
- homepage.icon=mdi-creation
|
|
- homepage.description=Mistral-Small-4-119B-2603 MoE (NVFP4) via vLLM (ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/mistral-small-4-heretic/compose.yaml
|
|
# mistral-small-4-heretic — abliterated Mistral Small 4 (heretic NVFP4) as a
|
|
# DROP-IN for the official mistral-small-4 backend.
|
|
#
|
|
# Serves darkc0de/Mistral-Small-4-119B-2603-heretic, quantized to NVFP4 in-house
|
|
# (vision tower kept bf16) and converted to Mistral native format. Built + validated
|
|
# 2026-06-17 — see tools/mistral-small4-nvfp4/ for the build pipeline.
|
|
#
|
|
# WHY a separate stack: GPU0 fits only one mistral-class model (~65-70 GB), so this
|
|
# is a backend SWAP, not a co-tenant. Bring it up only after downing the official
|
|
# mistral-small-4 stack. It serves under --served-model-name mistral-small-4 on the
|
|
# SAME port (8010), so litellm's mistral-small-4 + mistral-small-4-reasoning entries
|
|
# route here with NO litellm change. Revert = down this, `up -d` the official stack.
|
|
#
|
|
# DIFFERENCES vs the official compose (everything else mirrors it for a faithful
|
|
# drop-in — TP=1, util 0.93, MLA, reasoning + tool-call parsers, prefix caching):
|
|
# - model is a LOCAL native dir (mounted /model), not an HF id, so it needs the
|
|
# native loader flags: --config-format/--load-format/--tokenizer-mode mistral.
|
|
# - distinct container_name (vllm-mistral4-heretic) so it can be staged without
|
|
# colliding with the official container.
|
|
|
|
name: mistral-small-4-heretic
|
|
|
|
services:
|
|
vllm-mistral4-heretic:
|
|
image: ${MISTRAL_IMAGE}
|
|
container_name: ${MISTRAL_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${MISTRAL_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
# The in-house heretic native NVFP4 checkpoint (read-only).
|
|
- /tank/aimodels/quant-work/heretic-native-nvfp4:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- /model
|
|
# Native Mistral format (params.json + consolidated*.safetensors + tekken).
|
|
- --config-format
|
|
- mistral
|
|
- --load-format
|
|
- mistral
|
|
- --tokenizer-mode
|
|
- mistral
|
|
# SAME served name as the official → litellm routes here unchanged.
|
|
- --served-model-name
|
|
- mistral-small-4
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --tensor-parallel-size
|
|
- "1"
|
|
- --gpu-memory-utilization
|
|
- ${MISTRAL_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${MISTRAL_MAX_MODEL_LEN}
|
|
- --attention-backend
|
|
- TRITON_MLA
|
|
- --tool-call-parser
|
|
- mistral
|
|
- --enable-auto-tool-choice
|
|
- --reasoning-parser
|
|
- mistral
|
|
- --max-num-seqs
|
|
- ${MISTRAL_MAX_NUM_SEQS}
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${MISTRAL_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 600s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Mistral Small 4 (heretic NVFP4)
|
|
- homepage.icon=mdi-creation
|
|
- homepage.description=Abliterated Mistral-Small-4 (heretic NVFP4) drop-in via vLLM (ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/ms32-24b-angel/compose.yaml
|
|
# ms32-24b-angel — allura-org/MS3.2-24b-Angel (Mistral-Small-3.2-24B RP/fiction finetune) on ana-ml2 GPU 0.
|
|
# Serves the char-rp slot (:8016), REPLACING the qwen aeon-rp seat. NVFP4 (compressed-tensors: MLP quantized,
|
|
# vision tower + attention + lm_head kept bf16). NON-reasoning RP model — so NO MTP/spec-decode, NO GDN
|
|
# mamba-cache, NO reasoning-parser. Mistral tokenizer (--tokenizer-mode mistral, per the model card).
|
|
# Served under the aeon-rp names so the LiteLLM gateway char-rp / char-rp-reasoning routing stays transparent.
|
|
# Sampling defaults live at the gateway (RP: temp 1.2 / min_p 0.1 / rep 1.05, per the card + community).
|
|
# REVERT: `docker compose down` here + `docker compose up -d vllm-aeon-rp` in ../qwen36-27b-aeon.
|
|
|
|
name: ms32-24b-angel
|
|
|
|
services:
|
|
vllm-angel-rp:
|
|
image: ${ANGEL_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${ANGEL_CONTAINER_NAME:-vllm-angel-rp}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${ANGEL_PORT:-8016}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
- ${ANGEL_MODEL:-/tank/aimodels/ms32-24b-angel-nvfp4}:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
command:
|
|
- /model
|
|
- --served-model-name
|
|
- ${ANGEL_SERVED_NAME:-qwen3.6-27b-aeon-rp}
|
|
- ${ANGEL_SERVED_NAME_ALT:-qwen3.6-27b-aeon-rp-thinking}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --quantization
|
|
- compressed-tensors
|
|
- --tokenizer-mode
|
|
- ${ANGEL_TOKENIZER_MODE:-auto}
|
|
- --gpu-memory-utilization
|
|
- ${ANGEL_GPU_MEM_UTIL:-0.35}
|
|
- --max-model-len
|
|
- ${ANGEL_MAX_MODEL_LEN:-131072}
|
|
- --max-num-seqs
|
|
- ${ANGEL_MAX_NUM_SEQS:-4}
|
|
- --dtype
|
|
- auto
|
|
- --kv-cache-dtype
|
|
- ${ANGEL_KV_CACHE_DTYPE:-fp8}
|
|
- --enable-prefix-caching
|
|
- --limit-mm-per-prompt
|
|
- '{"image": 0}'
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${ANGEL_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=MS3.2-24B Angel (NVFP4, vision) — char-rp
|
|
- homepage.icon=mdi-drama-masks
|
|
- homepage.description=allura-org MS3.2-24B Angel, uncensored RP/fiction (char-rp), ana-ml2 GPU 0
|
|
- homepage.href=http://10.250.50.54:${ANGEL_PORT:-8016}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/parakeet/compose.yaml
|
|
services:
|
|
parakeet-stt:
|
|
image: parakeet-stt
|
|
ports:
|
|
- 8300:8000
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
restart: unless-stopped
|
|
volumes:
|
|
- parakeet_cache:/root/.cache
|
|
networks:
|
|
- tnet
|
|
env_file:
|
|
- .env
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Parakeet
|
|
- homepage.icon=mdi-talk
|
|
- homepage.description=Parakeet STT (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:8300
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
volumes:
|
|
parakeet_cache: null
|
|
|
|
>>> /opt/docker/compose/qwen3.5-122b/compose.yaml
|
|
# qwen3.5-122b — bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4 on ana-ml2 GPU 0,
|
|
# REPLACING mistral-small-4 (operator 2026-06-19: down the heretic, serve this as
|
|
# the new general/`gen` model). Abliterated Qwen3.5 MoE (256 experts, 10B active),
|
|
# NVFP4 (compressed-tensors), HF format.
|
|
#
|
|
# SERVING (per the repo's serving/): Qwen3.5 MoE is a MULTIMODAL arch but this
|
|
# checkpoint is text-only, so vLLM needs the repo's text-only PATCH applied before
|
|
# startup. We reuse the repo's entrypoint.sh (applies the patch, then runs vLLM) and
|
|
# vllm_patches/, mounted from the downloaded model dir — keeps patch+checkpoint
|
|
# version-coupled. Thinking split via litellm extra_body chat_template_kwargs
|
|
# (enable_thinking) + --reasoning-parser qwen3 (mirrors the qwen3.6-35b-a3b pattern).
|
|
#
|
|
# DISPLACEMENT: GPU 0 fits one mistral-class model; bring this up only after downing
|
|
# mistral-small-4-heretic. REVERT = down this, `up -d` the heretic stack.
|
|
#
|
|
# NOTE: image is vllm/vllm-openai:latest per the repo (the patch targets latest) —
|
|
# MUTABLE tag; pin a digest once a known-good version is established.
|
|
#
|
|
# Tunables in .env.
|
|
|
|
name: qwen3.5-122b
|
|
|
|
services:
|
|
vllm-qwen35-122b:
|
|
image: ${QWEN35_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${QWEN35_CONTAINER_NAME:-vllm-qwen35-122b}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QWEN35_PORT:-8013}:8000"
|
|
environment:
|
|
- ROLE=head
|
|
- TP_SIZE=1
|
|
- MODEL_CONTAINER_PATH=/models/qwen
|
|
- SERVED_MODEL_NAME=${QWEN35_SERVED_NAME:-qwen3.5-122-a10b}
|
|
- HOST_PORT=8000
|
|
- MAX_MODEL_LEN=${QWEN35_MAX_MODEL_LEN:-131072}
|
|
- MAX_NUM_SEQS=${QWEN35_MAX_NUM_SEQS:-8}
|
|
- GPU_MEMORY_UTILIZATION=${QWEN35_GPU_MEM_UTIL:-0.95}
|
|
- MAX_NUM_BATCHED_TOKENS=${QWEN35_MAX_NUM_BATCHED_TOKENS:-32768}
|
|
# --reasoning-parser qwen3 surfaces <think>…</think> as reasoning_content;
|
|
# the thinking on/off itself is per-request (litellm chat_template_kwargs).
|
|
# --enable-auto-tool-choice + --tool-call-parser: Qwen3.5 emits XML tool calls
|
|
# <tool_call><function=NAME><parameter=K>V</parameter></function></tool_call>
|
|
# (NOT Hermes JSON), so the parser is qwen3_xml. Without these flags vLLM never
|
|
# parses tool calls (tool-calling is broken). The bjk110 repo command omitted them.
|
|
- VLLM_EXTRA_ARGS=--reasoning-parser qwen3 --enable-chunked-prefill --enable-auto-tool-choice --tool-call-parser qwen3_xml
|
|
- NVIDIA_VISIBLE_DEVICES=0
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
volumes:
|
|
# checkpoint + the repo's entrypoint/patch (downloaded with the model).
|
|
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4:/models/qwen:ro
|
|
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/serving/entrypoint.sh:/entrypoint.sh:ro
|
|
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/vllm_patches:/patches:ro
|
|
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/.cache/vllm:/root/.cache/vllm
|
|
entrypoint: ["bash", "/entrypoint.sh"]
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QWEN35_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Qwen3.5-122B-A10B (abliterated NVFP4)
|
|
- homepage.icon=mdi-creation
|
|
- homepage.description=Abliterated Qwen3.5 122B-A10B NVFP4, the new `gen` model (ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${QWEN35_PORT:-8013}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/qwen35-vl/compose.yaml
|
|
# qwen35-vl — Qwen3.5-9B vision-language model (FP8) on ana-ml2.
|
|
#
|
|
# Co-located on GPU 1 with the granite summarizer + embed/rerank/reward trio
|
|
# (GPU 0 is deliberately kept free for hot-reloading large models). Serves on
|
|
# :8007, fronted by the LiteLLM gateway as `qwen3.5-9b-fp8`.
|
|
#
|
|
# WHY A PINNED NIGHTLY DIGEST (not :latest): vLLM :latest (v0.19.1) quantizes
|
|
# the Qwen3.5-VL *vision tower* under --quantization fp8, producing garbage
|
|
# vision output (the language model is unaffected — it answers text fine but
|
|
# "sees" noise). The nightly correctly excludes the vision tower, so vision
|
|
# works while the LM still gets the FP8 throughput/VRAM win. We pin the exact
|
|
# nightly digest for reproducibility — a moving :nightly tag would silently
|
|
# change the engine. WATCH: once the vision-FP8 exclusion lands in a stable
|
|
# release, re-pin to :latest and drop this note.
|
|
#
|
|
# WHY util 0.40 (not the trio's tiny values): the model needs ~34 GB just to
|
|
# start at 32k context (FP8 weights + BF16 vision tower + graph capture + 32k
|
|
# profiling). On shared GPU 1 (prod uses ~46 GB, ~48 GB free) this vLLM build
|
|
# requires free >= util*total, capping util at ~0.51 here; 0.40 (~38 GB) sits
|
|
# above the ~34 GB floor with ~10 GB card headroom.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
name: qwen35-vl
|
|
|
|
services:
|
|
vllm-qwen35:
|
|
image: ${QWEN_IMAGE}
|
|
container_name: ${QWEN_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QWEN_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${QWEN_MODEL}
|
|
- --served-model-name
|
|
- ${QWEN_SERVED_NAME}
|
|
- --quantization
|
|
- fp8
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${QWEN_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${QWEN_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
# Prefix caching pinned ON (the nightly defaults it OFF). Free win for the
|
|
# text-chat path; marginal for vision (each image is a distinct prefix).
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QWEN_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Qwen3.5-9B VL (FP8)
|
|
- homepage.icon=mdi-image-search
|
|
- homepage.description=Qwen3.5-9B vision-language (FP8) via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/qwen36-27b-aeon/compose.yaml
|
|
# qwen36-27b-aeon — AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored on ana-ml2 GPU 0,
|
|
# REPLACING qwopus3.5-122b as the `gen` model (operator 2026-07-05).
|
|
#
|
|
# Dense 27B, qwen3_5 GDN-hybrid arch (full-attn + Gated DeltaNet SSM) — same family
|
|
# as qwopus, VISION-INTACT (Qwen3_5ForConditionalGeneration, vision tower preserved
|
|
# at bf16), abliterated (abliterix v1.4, 0/100 refusals), native MTP head grafted,
|
|
# Apache-2.0, 131K default ctx. Served NVFP4 (ModelOpt) on Blackwell's FP4 cores.
|
|
#
|
|
# TWO CO-LOCATED INSTANCES on GPU 0 (operator wants both behaviours resident at once;
|
|
# MTP is a serve-time config, NOT per-request, so one endpoint can't do both):
|
|
# vllm-aeon-gen (:8015, served qwen3.6-27b-aeon) — MTP OFF, general/concurrent
|
|
# serve. Backs gateway gen / gen-reasoning / summarizer-large.
|
|
# vllm-aeon-rp (:8016, served qwen3.6-27b-aeon-rp) — native MTP ON (qwen3_5_mtp
|
|
# n=3), low-concurrency single-seat RP. Backs gateway char-rp.
|
|
# Why MTP off for the general serve: measured on qwopus, MTP helps single-stream
|
|
# (+12% N=1) but HURTS moderate concurrency (-15..-20% N=4) and silently drops
|
|
# min_p/logit_bias — wrong for a shared multi-consumer endpoint. Right only for a
|
|
# dedicated single-stream seat (the RP one). [[reference_gen_qwopus_122b]]
|
|
#
|
|
# VRAM budget (2 weight copies, no sharing): full NVFP4 = 27GB ea. gen util 0.45
|
|
# (~43GB) + rp util 0.40 (~38GB) = ~81GB / 96GB, ~15GB margin. depends_on:
|
|
# service_healthy sequences gen-first so the util reservation doesn't race → OOM.
|
|
# If margin bites at warmup (vision-encoder + big-vocab sampler warmup, cf. qwen36-vl),
|
|
# point AEON_RP_MODEL at the 21GB XS variant (frees ~6GB) via .env — no compose edit.
|
|
#
|
|
# NVFP4 is ModelOpt format → --quantization modelopt (vLLM also auto-detects; explicit
|
|
# is belt-and-suspenders). --mamba-cache-dtype float32 for the GDN/SSM state (AEON
|
|
# deploy guide + vLLM recipe Mamba-cache note). Tool-calling qwen3_coder + reasoning
|
|
# qwen3 (per the AEON card), same as qwopus.
|
|
#
|
|
# ⚠️ DEPLOYABILITY — load-test before trusting: multimodal + ModelOpt-NVFP4 on THIS
|
|
# brand-new arch, and native qwen3_5_mtp spec-decode, are unproven on our stock vLLM
|
|
# image. If stock can't serve it, the AEON patched image (ghcr.io/aeon-7/aeon-vllm-
|
|
# ultimate, PRs #41703/#40898) is the fallback — but that's really for DFlash; native
|
|
# MTP + base inference should ride stock >= 0.23.0. Set AEON_IMAGE in .env.
|
|
#
|
|
# REVERT: `docker compose down` here + `docker compose up -d` the qwopus3.5-122b stack
|
|
# (still staged) + revert the litellm gen/gen-reasoning/summarizer-large records.
|
|
# All tunables in .env — edit that, not this file.
|
|
|
|
name: qwen36-27b-aeon
|
|
|
|
services:
|
|
# ── General serve — MTP OFF, concurrent. gen / gen-reasoning / summarizer-large. ──
|
|
vllm-aeon-gen:
|
|
# Per-service image so gen can stay pinned to a known-good vLLM while rp tests a new one.
|
|
image: ${AEON_GEN_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${AEON_GEN_CONTAINER_NAME:-vllm-aeon-gen}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${AEON_GEN_PORT:-8015}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
- ${AEON_GEN_MODEL:-/tank/aimodels/qwen36-27b-aeon-nvfp4}:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
# Reclaims PyTorch reserved-but-unallocated fragmentation so the co-located
|
|
# util split doesn't strand VRAM (same knob qwopus needed for the MoE workspace).
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
command:
|
|
- /model
|
|
# TWO served-names: base + a `-thinking` alias. LiteLLM keys deployments by
|
|
# (model, api_base), so gen and gen-reasoning MUST use distinct model names or a
|
|
# thinking-off request mutates the shared litellm_params and clobbers the other's
|
|
# enable_thinking (the shared-config-mutation footgun). gen-reasoning routes to the
|
|
# `-thinking` name; gen/summarizer-large route to the base name.
|
|
- --served-model-name
|
|
- ${AEON_GEN_SERVED_NAME:-qwen3.6-27b-aeon}
|
|
- ${AEON_GEN_SERVED_NAME_THINK:-qwen3.6-27b-aeon-thinking}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --quantization
|
|
- ${AEON_GEN_QUANT:-modelopt}
|
|
- --gpu-memory-utilization
|
|
- ${AEON_GEN_GPU_MEM_UTIL:-0.45}
|
|
- --max-model-len
|
|
- ${AEON_GEN_MAX_MODEL_LEN:-131072}
|
|
# Keep concurrency modest: big-vocab sampler warmup allocates a large tensor
|
|
# (qwen36-vl OOM'd at the default 1024 on a shared GPU). 16 is ample here.
|
|
- --max-num-seqs
|
|
- ${AEON_GEN_MAX_NUM_SEQS:-16}
|
|
- --max-num-batched-tokens
|
|
- "16384"
|
|
- --trust-remote-code
|
|
- --dtype
|
|
- auto
|
|
# GDN/SSM (Gated DeltaNet) state cache — float32 per the AEON deploy guide.
|
|
- --mamba-cache-dtype
|
|
- float32
|
|
- --kv-cache-dtype
|
|
- ${AEON_GEN_KV_CACHE_DTYPE:-fp8}
|
|
- --enable-prefix-caching
|
|
- --enable-chunked-prefill
|
|
- --limit-mm-per-prompt
|
|
- '{"image": 4}'
|
|
- --reasoning-parser
|
|
- ${AEON_GEN_REASONING_PARSER:-qwen3}
|
|
- --enable-auto-tool-choice
|
|
- --tool-call-parser
|
|
- qwen3_coder
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${AEON_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Inference
|
|
- homepage.name=Qwen3.6-27B AEON (NVFP4, vision) — gen
|
|
- homepage.icon=mdi-creation
|
|
- homepage.description=Uncensored Qwen3.6-27B multimodal NVFP4, the `gen` model (ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${AEON_GEN_PORT:-8015}/docs
|
|
|
|
# ── RP seat — native MTP ON, low concurrency, single-seat. char-rp. ──
|
|
vllm-aeon-rp:
|
|
image: ${AEON_RP_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${AEON_RP_CONTAINER_NAME:-vllm-aeon-rp}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
# Sequence AFTER the general serve is healthy so the two util reservations on the
|
|
# shared GPU don't race into an OOM (gen reserves its 0.45 first, then rp its 0.40).
|
|
depends_on:
|
|
vllm-aeon-gen:
|
|
condition: service_healthy
|
|
ports:
|
|
- "${AEON_RP_PORT:-8016}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
# Point at the XS (21GB) variant via .env to buy ~6GB co-location margin.
|
|
- ${AEON_RP_MODEL:-/tank/aimodels/qwen36-27b-aeon-nvfp4}:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
command:
|
|
- /model
|
|
# base + `-thinking` alias (see the gen note): char-rp -> base, char-rp-reasoning -> -thinking.
|
|
- --served-model-name
|
|
- ${AEON_RP_SERVED_NAME:-qwen3.6-27b-aeon-rp}
|
|
- ${AEON_RP_SERVED_NAME_THINK:-qwen3.6-27b-aeon-rp-thinking}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --quantization
|
|
- ${AEON_RP_QUANT:-modelopt}
|
|
- --gpu-memory-utilization
|
|
- ${AEON_RP_GPU_MEM_UTIL:-0.40}
|
|
- --max-model-len
|
|
- ${AEON_RP_MAX_MODEL_LEN:-65536}
|
|
# Single-seat: low concurrency keeps warmup + KV small so it fits alongside gen.
|
|
- --max-num-seqs
|
|
- ${AEON_RP_MAX_NUM_SEQS:-2}
|
|
- --trust-remote-code
|
|
- --dtype
|
|
- auto
|
|
- --mamba-cache-dtype
|
|
- float32
|
|
- --kv-cache-dtype
|
|
- ${AEON_RP_KV_CACHE_DTYPE:-fp8}
|
|
- --enable-prefix-caching
|
|
- --limit-mm-per-prompt
|
|
- '{"image": 4}'
|
|
- --reasoning-parser
|
|
- ${AEON_RP_REASONING_PARSER:-qwen3}
|
|
# Native MTP speculative decode (the grafted head). n=3 per the AEON card's
|
|
# measured accept length (~3.3/3). MTP is why this seat exists separately.
|
|
- --speculative-config
|
|
- '{"method": "${AEON_RP_SPEC_METHOD:-qwen3_5_mtp}", "num_speculative_tokens": ${AEON_RP_SPEC_TOKENS:-3}}'
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${AEON_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Dormant
|
|
- homepage.name=Qwen3.6-27B AEON RP (NVFP4 + MTP) — char-rp
|
|
- homepage.icon=mdi-drama-masks
|
|
- homepage.description=Uncensored Qwen3.6-27B, native MTP single-seat RP (char-rp), ana-ml2 GPU 0
|
|
- homepage.href=http://10.250.50.54:${AEON_RP_PORT:-8016}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/qwen36-vl/compose.yaml
|
|
# qwen36-vl — Qwen3.6-35B-A3B vision-language MoE (official NVFP4) on ana-ml2.
|
|
#
|
|
# Replaces the qwen35-vl stack (Qwen3.5-9B) 2026-06-14. Co-located on GPU 1 with
|
|
# the granite summarizer + embed/rerank/reward trio. Serves on :8007.
|
|
#
|
|
# WHY NVFP4 now (swapped FROM FP8 2026-06-15): the nvidia ModelOpt NVFP4 MoE that
|
|
# was BROKEN on vLLM 0.19.1/0.22.0 (#44081, lm_head.input_scale) loads clean on
|
|
# 0.23.0 — the ModelOpt lm_head fix landed. So we cut FP8→NVFP4: ~20.4 GiB weights
|
|
# vs FP8's ~34 GiB (~40% lighter, ~13 GB reclaimed on GPU 1), faster single-stream
|
|
# on Blackwell's FP4 tensor cores, and the freed room funds fp16 KV + a granite
|
|
# context restore (see the rebalance note below). The NVFP4 checkpoint preserves
|
|
# the vision tower (ModelOpt leaves it high-precision) — VALIDATED by comfy-dev's
|
|
# real anatomy-judge A/B on 16 prod images: PASS, holds the load-bearing
|
|
# discrimination (gross-deformity reject + clean-pass), only shuffles already-
|
|
# unreliable sub-ceiling borderline-hand calls. brokkr's text/speed arm: parity
|
|
# except a minor multi-step chained-numeric-reasoning slip (W4A4 tell) — doesn't
|
|
# bite the vision-judge role; flag for any gateway consumer doing chained math.
|
|
# Requires vLLM >= 0.23.0 (pinned by digest in .env). NO --quantization flag
|
|
# (vLLM auto-detects the checkpoint's NVFP4).
|
|
#
|
|
# NAMING: served ONLY as its TRUE name `qwen3.6-35b-a3b`. A model is never aliased
|
|
# under a prior model's name — a caller asking for `qwen3.5-9b-fp8` (a 9B dense)
|
|
# must NOT be silently handed this 35B-A3B MoE; that's a downstream-confusion
|
|
# footgun. The legacy `qwen3.5-9b-fp8` name is RETIRED. Consumers (Arbo's vision
|
|
# hero-judge, stacks/arbo v0.11.3+) migrate to `qwen3.6-35b-a3b` — they 404 on the
|
|
# old name until they repoint, which is the correct loud signal (notified 2026-06-14).
|
|
#
|
|
# GPU-1 REBALANCE (2026-06-15, pinned): the NVFP4 swap freed ~13 GB, redistributed —
|
|
# qwen36 NVFP4: util 0.46→0.32 (~31 GB: 20.4 GB weights + fp16 KV + graph).
|
|
# fp16 KV (we DROPPED --kv-cache-dtype fp8) — the freed room buys back full-
|
|
# precision KV; hybrid attn (10/40 full-attn) keeps even fp16 KV affordable.
|
|
# granite: RESTORED 0.24→0.34, max-len 65536→131072 (gives back the context
|
|
# sacrificed for FP8 qwen — the FP8-vs-maxed-granite tradeoff is now undone).
|
|
# trio (embed/rerank/reward) unchanged at floor.
|
|
# Total GPU-1 util ~0.82 → ~17 GB headroom (was a tight ~5 GB).
|
|
#
|
|
# THINKING TOGGLE: this is ONE hybrid checkpoint (not separate Instruct/Thinking
|
|
# downloads) with a Qwen3-style per-request `enable_thinking` switch. The chat
|
|
# template defaults thinking ON (`<think>\n`); passing
|
|
# `chat_template_kwargs={"enable_thinking":false}` emits the empty
|
|
# `<think>\n\n</think>\n\n` block (no reasoning). We run --reasoning-parser qwen3
|
|
# (model-matched — its vLLM docstring describes THIS checkpoint) so ONE endpoint
|
|
# serves BOTH modes cleanly: thinking-ON splits <think>…</think> into
|
|
# reasoning_content; thinking-OFF routes everything to content. The gateway
|
|
# selects the mode per model_name (stacks/litellm/conf/config.yaml):
|
|
# qwen3.6-35b-a3b → enable_thinking:false (non-thinking DEFAULT)
|
|
# qwen3.6-35b-a3b-thinking → enable_thinking:true (opt-in reasoning)
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
|
|
name: qwen36-vl
|
|
|
|
services:
|
|
vllm-qwen36:
|
|
image: ${QWEN_IMAGE}
|
|
container_name: ${QWEN_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QWEN_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${QWEN_MODEL}
|
|
# Pre-quantized NVFP4 (ModelOpt) checkpoint → NO --quantization (vLLM auto-
|
|
# detects; the vision tower is left high-precision by the producer).
|
|
- --served-model-name
|
|
- qwen3.6-35b-a3b
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${QWEN_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${QWEN_MAX_MODEL_LEN}
|
|
# Cap concurrency: vLLM warms the sampler with max_num_seqs dummy requests,
|
|
# and this model's 248K vocab makes that warmup tensor huge — the default
|
|
# 1024 OOMs on a shared GPU even though weights+KV fit. 32 is ample for a
|
|
# vision endpoint (the summarizer carries the concurrency, not this).
|
|
- --max-num-seqs
|
|
- ${QWEN_MAX_NUM_SEQS}
|
|
# fp16 KV (no --kv-cache-dtype): the NVFP4 swap freed enough room to run
|
|
# full-precision KV — better than the fp8 KV the FP8 build needed to fit.
|
|
- --trust-remote-code
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
# Model-matched reasoning parser for the hybrid thinking toggle (see header).
|
|
# Splits <think>…</think> into reasoning_content when thinking is ON; routes
|
|
# all output to content when the empty think-block signals thinking OFF — so
|
|
# this single :8007 endpoint serves both the non-thinking default and the
|
|
# qwen3.6-35b-a3b-thinking gateway variant.
|
|
- --reasoning-parser
|
|
- qwen3
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QWEN_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 300s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=Qwen3.6-35B-A3B VL (NVFP4)
|
|
- homepage.icon=mdi-image-search
|
|
- homepage.description=Qwen3.6-35B-A3B vision-language MoE (NVFP4) via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/qwen-image-bench/compose.yaml
|
|
# qwen-image-bench — flukethoughts/Qwen-Image-Bench-NVFP4 on ana-ml2 GPU 1,
|
|
# REPLACING qwen3.6-35b-a3b (operator 2026-06-19). Qwen's text-to-image quality
|
|
# JUDGE model (vision-language, NVFP4 weights / vision tower bf16). NOT generative —
|
|
# it scores T2I outputs on 5 dims (overall quality, prompt match, aesthetic, LoRA
|
|
# activation, confidence).
|
|
#
|
|
# Arch: Qwen3_5ForConditionalGeneration (dense Qwen3.5 hybrid SSM+attn + vision),
|
|
# ~17B / ~20GB NVFP4. VISION-INTACT → served as multimodal; NO text-only patch
|
|
# (unlike the qwen3.5-122b gen model, which had text-only weights). vLLM
|
|
# production-validated per the model card.
|
|
#
|
|
# ⚠️ qwen3.6-35b-a3b was arbo's hero-judge (comfy-dev consumer). Downing it breaks
|
|
# arbo's judging until comfy-dev repoints to qwen-image-bench (different I/O — a
|
|
# 5-dim verdict vs a general VL judge). comfy-dev notified.
|
|
#
|
|
# DISPLACEMENT: GPU 1 is shared (granite/selene/embed/rerank/reward). qwen3.6 used
|
|
# util 0.34 (~33GB); down it first, then this fits at util ~0.22 (~21GB). REVERT =
|
|
# down this, `up -d` the qwen36 stack.
|
|
#
|
|
# Tunables in .env.
|
|
|
|
name: qwen-image-bench
|
|
|
|
services:
|
|
vllm-qwen-image-bench:
|
|
image: ${QIB_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${QIB_CONTAINER_NAME:-vllm-qwen-image-bench}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QIB_PORT:-8014}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
- /tank/aimodels/qwen-image-bench-nvfp4:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- /model
|
|
- --served-model-name
|
|
- qwen-image-bench
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${QIB_GPU_MEM_UTIL:-0.32}
|
|
- --max-model-len
|
|
- ${QIB_MAX_MODEL_LEN:-32768}
|
|
- --max-num-seqs
|
|
- ${QIB_MAX_NUM_SEQS:-8}
|
|
- --trust-remote-code
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QIB_GPU_ID:-1}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 600s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=Qwen-Image-Bench (T2I judge, NVFP4)
|
|
- homepage.icon=mdi-image-check
|
|
- homepage.description=Qwen text-to-image quality judge (NVFP4, vision-intact) on ana-ml2 GPU 1
|
|
- homepage.href=http://10.250.50.54:${QIB_PORT:-8014}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/qwopus3.5-122b/compose.yaml
|
|
# qwopus3.5-122b — OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4
|
|
# on ana-ml2 GPU 0, REPLACING the bjk110 text-only qwen3.5-122b as the `gen` model
|
|
# (operator 2026-06-19: "already ablated, already quanted, vision tower intact").
|
|
#
|
|
# Qwen3.5-122B-A10B MoE, Kimi-K2.6-distilled + abliterated, NVFP4 — and crucially
|
|
# VISION-INTACT (Qwen3_5MoeForConditionalGeneration + vision_config). So it serves as
|
|
# plain MULTIMODAL (no text-only patch, unlike the bjk110 checkpoint which had its
|
|
# vision weights stripped). vLLM carries the arch natively.
|
|
#
|
|
# Served under --served-model-name qwen3.5-122-a10b so the existing litellm records
|
|
# (gen / gen-reasoning / qwen3.5-122-a10b[-reasoning] / qwen-large[-reasoning]) route
|
|
# here UNCHANGED — the operator's "replace those records with this model". The thinking
|
|
# split (chat_template_kwargs.enable_thinking) + tool-calling (qwen3_coder — the
|
|
# OpenYourMind card's specified parser for this checkpoint's XML tool calls).
|
|
#
|
|
# REVERT: down this; the bjk110 qwen3.5-122b stack is still staged.
|
|
# Tunables in .env.
|
|
|
|
name: qwopus3.5-122b
|
|
|
|
services:
|
|
vllm-qwopus35-122b:
|
|
image: ${QWOPUS_IMAGE:-vllm/vllm-openai:latest}
|
|
container_name: ${QWOPUS_CONTAINER_NAME:-vllm-qwopus35-122b}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${QWOPUS_PORT:-8013}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
- /tank/aimodels/qwopus3.5-122b-nvfp4:/model:ro
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
# Reclaims PyTorch's reserved-but-unallocated fragmentation (4.2GB was stranded at
|
|
# util 0.96, starving the FusedMoE workspace → OOM by 0.1GB). Lets the 3.09GB MoE
|
|
# workspace allocate cleanly. Same knob the bjk110 qwen3.5-122b stack ran.
|
|
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
|
command:
|
|
- /model
|
|
- --served-model-name
|
|
- ${QWOPUS_SERVED_NAME:-qwen3.5-122-a10b}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${QWOPUS_GPU_MEM_UTIL:-0.92}
|
|
- --max-model-len
|
|
- ${QWOPUS_MAX_MODEL_LEN:-131072}
|
|
- --max-num-seqs
|
|
- ${QWOPUS_MAX_NUM_SEQS:-8}
|
|
- --max-num-batched-tokens
|
|
- "32768"
|
|
- --trust-remote-code
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
- --enable-chunked-prefill
|
|
# FULL 256K context on the STABLE image. fp8 KV (near-lossless) measured an 11.8GB
|
|
# pool = 934,600 tokens = 3.5x concurrency at the full 262144 window. CUDA graphs ON
|
|
# (no --enforce-eager) for decode tok/s. BINDING LIMIT = the FusedMoE transient
|
|
# workspace (3.09GB, allocated OUTSIDE vLLM's budget into free VRAM): at util 0.96
|
|
# only 2.99GB was free → OOM by 0.1GB, worsened by 4.2GB PyTorch fragmentation.
|
|
# FIX = expandable_segments (env above, reclaims the fragmentation) + util 0.95 for
|
|
# margin. The card can't go to 0 free — this workspace is the floor. video kept
|
|
# ENABLED (operator wants it; banked at util 0.95 with headroom) — the video encoder
|
|
# profiling eats into the budget so KV concurrency drops some, but stays well above 2x.
|
|
- --kv-cache-dtype
|
|
- ${QWOPUS_KV_CACHE_DTYPE:-fp8}
|
|
- --limit-mm-per-prompt
|
|
- '{"image": 2, "video": 1}'
|
|
# reasoning split + tool-calling. The OpenYourMind card specifies qwen3_coder
|
|
# as the tool-call parser for this checkpoint (Qwen3.5 XML tool-call format).
|
|
- --reasoning-parser
|
|
- qwen3
|
|
- --enable-auto-tool-choice
|
|
- --tool-call-parser
|
|
- qwen3_coder
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${QWOPUS_GPU_ID:-0}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 900s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Dormant
|
|
- homepage.name=Qwopus3.5-122B-A10B (abliterated NVFP4, vision)
|
|
- homepage.icon=mdi-creation
|
|
- homepage.description=Kimi-distilled abliterated Qwen3.5-122B-A10B NVFP4, vision-intact, the `gen` model (ana-ml2 GPU 0)
|
|
- homepage.href=http://10.250.50.54:${QWOPUS_PORT:-8013}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/selene/compose.yaml
|
|
# selene — AtlaAI Selene 1 Mini (Llama 3.1 8B) judge/eval model on ana-ml2 GPU 1.
|
|
#
|
|
# Restores the judge that went offline when llama-swap was downed (it was the
|
|
# Q6_K GGUF `selene-1-mini-8b` in the llama-swap zoo). Re-served on vLLM at the
|
|
# operator's request, FP8 (NVFP4 had no pre-made checkpoint and W4A4 is too
|
|
# aggressive for a precision judge validated at Q6_K — FP8 ≥ Q6_K fidelity).
|
|
#
|
|
# FP8 = vLLM DYNAMIC --quantization fp8 (W8A8) of the bf16 AtlaAI checkpoint —
|
|
# no offline quant needed, near-lossless, and Selene is text-only Llama 3.1 so
|
|
# there's NO vision tower for dynamic fp8 to noise-quantize (the qwen35-VL
|
|
# footgun doesn't apply here). ~8 GiB weights on GPU 1's headroom.
|
|
#
|
|
# Co-tenant on GPU 1 with qwen36 (NVFP4) + granite + embed/rerank/reward. Sized
|
|
# to fit the ~24 GB headroom while leaving GPU 1 a safe buffer (see .env).
|
|
# Served ONLY as `selene-1-mini-8b` (the name its consumers know). All tunables
|
|
# live in .env.
|
|
|
|
name: selene
|
|
|
|
services:
|
|
vllm-selene:
|
|
image: ${SELENE_IMAGE}
|
|
container_name: ${SELENE_CONTAINER_NAME}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${SELENE_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${SELENE_MODEL}
|
|
# Dynamic FP8 (W8A8) from the bf16 checkpoint — no pre-quant needed.
|
|
- --quantization
|
|
- fp8
|
|
- --served-model-name
|
|
- selene-1-mini-8b
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${SELENE_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${SELENE_MAX_MODEL_LEN}
|
|
- --max-num-seqs
|
|
- ${SELENE_MAX_NUM_SEQS}
|
|
# fp8 KV — matches the judge's old q8 KV posture + keeps the pool compact
|
|
# on the shared card.
|
|
- --kv-cache-dtype
|
|
- fp8
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${SELENE_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=Selene 1 Mini 8B (judge, FP8)
|
|
- homepage.icon=mdi-gavel
|
|
- homepage.description=AtlaAI Selene 1 Mini Llama-3.1-8B judge (FP8) via vLLM (ana-ml2 GPU1)
|
|
- homepage.href=http://10.250.50.54:${SELENE_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/vibevoice/compose.yaml
|
|
services:
|
|
vibevoice:
|
|
container_name: vibevoice
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
count: all
|
|
capabilities:
|
|
- gpu
|
|
ports:
|
|
- 8745:8745
|
|
volumes:
|
|
- /tank/vibevoice/hf:/root/.cache/huggingface
|
|
- /tank/vibevoice/voices:/app/voices
|
|
- /tank/vibevoice/state:/var/lib/eworker
|
|
environment:
|
|
- ENABLE_1_5B=true
|
|
- ENABLE_LARGE=true
|
|
- AUTH_REQUIRED=true
|
|
- CORS_ENABLED=true
|
|
- ALLOWED_ORIGINS=*
|
|
image: eworkerinc/vibevoice:latest
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI Systems
|
|
- homepage.name=VibeVoice
|
|
- homepage.icon=mdi-chat
|
|
- homepage.description=EWorkerStudio VibeVoice
|
|
- homepage.href=http://10.250.50.54:8745
|
|
restart: unless-stopped
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
>>> /opt/docker/compose/vllm/compose.yaml
|
|
# vLLM — Qwen3 Embedding + Reranker + Skywork Reward-V2 classifier.
|
|
#
|
|
# Originally created to replace the unmaintained Infinity stack (embed +
|
|
# rerank); generalized 2026-05-13 to host any vLLM-served model on ana-ml2,
|
|
# starting with the Skywork-Reward-V2-Llama-3.1-8B reward classifier
|
|
# (AWQ-quantized locally, served from /tank/aimodels/llm/).
|
|
#
|
|
# vLLM runs one model per process, so this stack brings up three containers
|
|
# sharing a single GPU:
|
|
#
|
|
# vllm-embed — Qwen3-Embedding served as an OpenAI /v1/embeddings server
|
|
# vllm-rerank — Qwen3-Reranker served as a /rerank + /score server
|
|
# vllm-reward — Skywork-Reward-V2-Llama-3.1-8B-AWQ served as a /classify scorer
|
|
#
|
|
# The reranker is a causal-LM checkpoint; --hf-overrides re-maps it to
|
|
# Qwen3ForSequenceClassification so vLLM's reranking endpoints work and the
|
|
# model only emits two class logits (no/yes) instead of the full 151k vocab.
|
|
#
|
|
# All tunables live in .env — edit that, not this file.
|
|
#
|
|
# Pre-download models to avoid first-run delay:
|
|
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
|
|
# --var hf_repo=Qwen/Qwen3-Embedding-0.6B
|
|
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
|
|
# --var hf_repo=Qwen/Qwen3-Reranker-0.6B
|
|
#
|
|
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
|
|
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on ana-ml2 and is
|
|
# bind-mounted into the reward service at /local-models. Not from HF Hub.
|
|
|
|
services:
|
|
vllm-embed:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-embed
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${EMBED_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${EMBED_MODEL}
|
|
- --served-model-name
|
|
- ${EMBED_MODEL}
|
|
- --runner
|
|
- pooling
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${EMBED_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${EMBED_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=vLLM Embed (Qwen3)
|
|
- homepage.icon=mdi-vector-arrange-below
|
|
- homepage.description=Qwen3 Embedding via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${EMBED_PORT}/docs
|
|
|
|
vllm-rerank:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-rerank
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${RERANK_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- ${RERANK_MODEL}
|
|
- --served-model-name
|
|
- ${RERANK_MODEL}
|
|
- --runner
|
|
- pooling
|
|
- --hf-overrides
|
|
- '{"architectures":["Qwen3ForSequenceClassification"],"classifier_from_token":["no","yes"],"is_original_qwen3_reranker":true}'
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${RERANK_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${RERANK_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=vLLM Rerank (Qwen3)
|
|
- homepage.icon=mdi-sort-variant
|
|
- homepage.description=Qwen3 Reranker via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${RERANK_PORT}/docs
|
|
|
|
vllm-reward:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-reward
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${REWARD_PORT}:8000"
|
|
volumes:
|
|
# AWQ output lives in the legacy llama-swap models tree, not the HF cache
|
|
# — bind-mount the LLM models dir read-only so the reward service can
|
|
# load it as a local-path HF-format model.
|
|
- /tank/aimodels/llm:/local-models:ro
|
|
environment:
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
- /local-models/Skywork-Reward-V2-Llama-3.1-8B-AWQ
|
|
- --served-model-name
|
|
- Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
|
|
# vLLM 0.19.1 deprecated --task in favor of --runner. The model's
|
|
# config.json declares `LlamaForSequenceClassification` so the
|
|
# pooling runner uses it as a classifier (single-label reward score)
|
|
# without needing an explicit task flag.
|
|
- --runner
|
|
- pooling
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${REWARD_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${REWARD_MAX_MODEL_LEN}
|
|
- --dtype
|
|
- auto
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 240s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Eval & Retrieval
|
|
- homepage.name=vLLM Reward (Skywork)
|
|
- homepage.icon=mdi-scale-balance
|
|
- homepage.description=Skywork-Reward-V2 8B classifier via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${REWARD_PORT}/docs
|
|
|
|
# Phi-4-mini (FP8) — summarizer + "dreaming" agent. Supersedes the
|
|
# llama-swap granite-4-small pin. Generative chat model (OpenAI
|
|
# /v1/chat/completions), so NO --runner pooling. FP8 on RTX 6000 Ada
|
|
# (cc 8.9): near-lossless, ~1.2x, ~6 GB.
|
|
vllm-granite:
|
|
image: vllm/vllm-openai:${VLLM_VERSION}
|
|
container_name: vllm-granite
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${GRANITE_PORT}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
environment:
|
|
- HF_HOME=/hfcache
|
|
- HF_HUB_CACHE=/hfcache/hub
|
|
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
command:
|
|
# Production summarizer (replaced phi4-mini 2026-06-05). Default = official
|
|
# IBM pre-quantized FP8 (compressed-tensors), loaded directly; FP8 is native
|
|
# on the RTX 6000 Ada (cc 8.9). Fallback to vLLM-native dynamic FP8 from
|
|
# BF16: GRANITE_MODEL=ibm-granite/granite-4.1-8b + GRANITE_QUANT=fp8.
|
|
- ${GRANITE_MODEL}
|
|
- --served-model-name
|
|
- ${GRANITE_SERVED_NAME}
|
|
- --quantization
|
|
- ${GRANITE_QUANT}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- ${GRANITE_GPU_MEM_UTIL}
|
|
- --max-model-len
|
|
- ${GRANITE_MAX_MODEL_LEN}
|
|
- --max-num-seqs
|
|
- ${GRANITE_MAX_NUM_SEQS}
|
|
- --dtype
|
|
- auto
|
|
# CUDA graphs ENABLED (no --enforce-eager) for decode throughput. Made
|
|
# room 2026-06-05 by right-sizing the embed/rerank/reward trio's KV pools
|
|
# (they were over-provisioned at 5.9x/2.0x/3.9x concurrency); GPU 1 now has
|
|
# ~17 GB free after granite, so graph-capture buffers fit. If the trio
|
|
# ever grows back, granite may need --enforce-eager again on this card.
|
|
# FP8 KV cache — halves KV memory; near-lossless on Ada (cc 8.9).
|
|
- --kv-cache-dtype
|
|
- ${GRANITE_KV_CACHE_DTYPE}
|
|
# Prefix caching pinned EXPLICIT (vLLM v1 defaults it on, but pin so a
|
|
# version flip can't silently disable it). Benched 2026-06-13: ~6.5x faster
|
|
# TTFT (45ms vs 292ms) on a shared ~4.5k-token summarizer template; soft/
|
|
# evictable KV, neutral when prefixes don't repeat — pure win for granite.
|
|
- --enable-prefix-caching
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${GRANITE_GPU_ID}"
|
|
capabilities:
|
|
- gpu
|
|
healthcheck:
|
|
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
|
interval: 30s
|
|
timeout: 10s
|
|
retries: 3
|
|
start_period: 180s
|
|
networks:
|
|
- tnet
|
|
labels:
|
|
- homepage.group=AI - Inference
|
|
- homepage.name=vLLM Granite 4.1 8B (summarizer)
|
|
- homepage.icon=mdi-text-box-outline
|
|
- homepage.description=Granite 4.1 8B FP8 via vLLM (ana-ml2)
|
|
- homepage.href=http://10.250.50.54:${GRANITE_PORT}/docs
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|
|
|
|
===== CONFIG LAYOUT (/opt/docker/conf/ — top 200 entries) =====
|
|
|
|
/opt/docker/conf
|
|
/opt/docker/conf/llama-swap
|
|
/opt/docker/conf/llama-swap/config.yaml
|
|
/opt/docker/conf/vllm
|
|
|
|
===== LISTENING PORTS =====
|
|
|
|
0.0.0.0:111
|
|
0.0.0.0:22
|
|
0.0.0.0:5001
|
|
0.0.0.0:7007
|
|
0.0.0.0:8001
|
|
0.0.0.0:8002
|
|
0.0.0.0:8003
|
|
0.0.0.0:8004
|
|
0.0.0.0:8011
|
|
0.0.0.0:8015
|
|
0.0.0.0:8016
|
|
0.0.0.0:8018
|
|
[::]:111
|
|
[::]:22
|
|
*:2375
|
|
[::]:5001
|
|
[::]:8001
|
|
[::]:8002
|
|
[::]:8003
|
|
[::]:8004
|
|
[::]:8011
|
|
[::]:8015
|
|
[::]:8016
|
|
[::]:8018
|
|
|
|
===== MODEL / HUGGINGFACE CACHES =====
|
|
|
|
/tank/aimodels/huggingface (604G)
|
|
hub entries:
|
|
CACHEDIR.TAG
|
|
datasets--HuggingFaceH4--ultrachat_200k
|
|
datasets--mlabonne--harmful_behaviors
|
|
datasets--mlabonne--harmless_alpaca
|
|
datasets--neuralmagic--calibration
|
|
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
|
|
datasets--wikitext
|
|
models--AtlaAI--Selene-1-Mini-Llama-3.1-8B
|
|
models--AxionML--Qwen3.5-9B-NVFP4
|
|
models--bartowski--Meta-Llama-3.1-8B-Instruct-GGUF
|
|
models--bartowski--NousResearch_Hermes-4-14B-GGUF
|
|
models--bartowski--TheDrummer_GLM-Steam-106B-A12B-v1-GGUF
|
|
models--bartowski--TheDrummer_Skyfall-31B-v4-GGUF
|
|
models--BeaverAI--Artemis-31B-v1i-GGUF
|
|
models--BeaverAI--Skyfall-R1-31B-v4a-GGUF
|
|
models--DavidAU--Qwen3.6-27B-Heretic2-Uncensored-Finetune-Thinking
|
|
models--drawais--Granite-4.1-30B-NVFP4
|
|
models--ibm-granite--granite-4.0-h-small-GGUF
|
|
models--ibm-granite--granite-4.0-h-tiny-GGUF
|
|
models--ibm-granite--granite-4.0-micro-GGUF
|
|
models--ibm-granite--granite-4.1-8b
|
|
models--ibm-granite--granite-4.1-8b-fp8
|
|
models--llmfan46--Qwen3.6-35B-A3B-uncensored-heretic-GGUF
|
|
models--microsoft--Phi-4-mini-instruct
|
|
models--mistralai--Mistral-Small-4-119B-2603-NVFP4
|
|
models--mradermacher--Daredevil-8B-abliterated-dpomix-GGUF
|
|
models--mradermacher--Qwen3-30B-A3B-abliterated-erotic-i1-GGUF
|
|
models--mradermacher--Qwen3.6-35B-A3B-abliterated-i1-GGUF
|
|
models--mradermacher--Selene-1-Mini-Llama-3.1-8B-i1-GGUF
|
|
models--murilonwt--granite-4.1-8b-NVFP4
|
|
|
|
/tank/aimodels/llm (1.1T)
|
|
|
|
/home/lkraven/.cache/huggingface (18G)
|
|
hub entries:
|
|
CACHEDIR.TAG
|
|
datasets--HuggingFaceH4--ultrachat_200k
|
|
datasets--Salesforce--wikitext
|
|
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
|
|
datasets--wikitext
|
|
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
|
|
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP
|
|
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
|
|
models--Astralyra--bge-reranker-large-Q8_0-GGUF
|
|
models--bicro--qwen3.5-abliterated-vision-merged
|
|
models--bjk110--Qwen3.5-122B-A10B-abliterated-NVFP4
|
|
models--darkc0de--Mistral-Small-4-119B-2603-heretic
|
|
models--DavidAU--Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
|
|
models--flukethoughts--Qwen-Image-Bench-NVFP4
|
|
models--ggml-org--embeddinggemma-300M-GGUF
|
|
models--ggml-org--Qwen3-Reranker-0.6B-Q8_0-GGUF
|
|
models--huihui-ai--Huihui-Qwen3.5-122B-A10B-abliterated
|
|
models--jinaai--jina-reranker-v3-GGUF
|
|
models--klnstpr--bge-reranker-v2-m3-Q8_0-GGUF
|
|
models--minhtd14--jina-reranker-v2-base-multilingual-Q8_0-GGUF
|
|
models--mistralai--Mistral-Medium-3.5-128B-EAGLE
|
|
models--mistralai--Mistral-Small-4-119B-2603-NVFP4
|
|
models--Mungert--Qwen3-Reranker-0.6B-GGUF
|
|
models--OpenYourMind--Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4
|
|
models--OpenYourMind--Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated
|
|
models--Qwen--Qwen3-Embedding-0.6B-GGUF
|
|
models--Qwen--Qwen3-Reranker-0.6B
|
|
models--RecViking--Mistral-Medium-3.5-128B-NVFP4
|
|
models--robbatt--Qwen3.6-40B-Deckard-NVFP4
|
|
models--Skywork--Skywork-Reward-V2-Llama-3.1-8B
|
|
|
|
|
|
===== DOCKER-ADJACENT SYSTEMD SERVICES =====
|
|
|
|
containerd.service running
|
|
docker.service running
|
|
nvidia-persistenced.service running
|
|
|
|
===== DONE =====
|
|
|
|
Paste the above back into the chat, or pass a path as argv[1] to save.
|