Files
esh-pfi-infrastructure/servers/fv-ml1/system-details.txt
T
vh 91bda3c480 fv-ml1: complete the cutover — rename, renumber, DNS, and the LiteLLM repoint
The box is physically at Fountain Valley, renamed, renumbered onto 10.251/16,
and serving inference again. This lands the repo half of that.

Host: hostname ana-ml2 -> fv-ml1, pinned to 10.251.50.54 by a dnsmasq
reservation so the address the runbook, DNS and LiteLLM all assume is the
address it actually has. Its headscale node is renamed too.

The sweep ran from scripts/fv-ml1-rename-sweep.sh, whose allowlist is the
reason this diff touches current-state files and not the record. Dated
persistent-memory entries, archival-memory and incident notes still say
ana-ml2 in 31 and 62 places respectively, because that is what the box was
when those things happened. Rewriting them would make the history lie.

LiteLLM was the load-bearing piece and needed more than the api_base sed the
runbook describes. Twenty api_base entries repointed, but a grep-and-verify
pass also caught a LIVE pass_through_endpoints target for the scalar-judge
reward route still on the old address -- an api_base-only substitution would
have left it dead. Four prose references describing current state were
repointed as well; one historical note recording where a hand-test was run
is deliberately left pointing at 10.250.50.54.

Two facts in the server tables were wrong and are corrected here. The site is
Fountain Valley, not Anaheim. And the box has FOUR RTX PRO 6000 Blackwell
Max-Q, not two -- verified by nvidia-smi -L and independently by PCI
enumeration of four GB202GL devices. That is 391 GB of VRAM rather than 196,
which changes what fits on it.

DNS: fv-ml1, fv-ml1-bmc and fv-gw added under the fv site via the piggyback
approach, scriberr re-homed, and the ana-ml2 records removed. Applied to all
three resolvers. The BMC record carries a warning that its 802.1q VLAN tag
must stay disabled -- it shipped tagging VLAN 250 into an untagged port,
which made it invisible to every network-side diagnostic and is the reason
it appeared dead through several cable changes.

Verified end to end: summarizer and sec both answer through the Anaheim
gateway across the mesh to FV seats on different ports.
2026-09-12 22:00:50 -07:00

2494 lines
100 KiB
Plaintext

===== HOST =====
Hostname: ana-ml2
Date: 2026-07-22T15:24:06-07:00
Uptime: up 5 weeks, 5 days, 1 hour, 55 minutes
OS: Debian GNU/Linux 13 (trixie)
Kernel: 6.12.74+deb13+1-amd64
Arch: x86_64
===== HARDWARE =====
CPU cores: 96
CPU model: AMD EPYC 9254 24-Core Processor
MemTotal: 566.6 GB
MemAvailable: 229.0 GB
===== GPUS =====
index, name, memory.total [MiB], memory.free [MiB], driver_version
0, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 9494 MiB, 580.65.06
1, NVIDIA RTX PRO 6000 Blackwell Max-Q Workstation Edition, 97887 MiB, 6238 MiB, 580.65.06
===== FILESYSTEMS (df) =====
Filesystem Size Used Avail Use% Mounted on
zroot/ROOT/debian 384G 271G 113G 71% /
efivarfs 128K 67K 57K 55% /sys/firmware/efi/efivars
/dev/sdb1 511M 92M 420M 18% /boot/efi
tank 8.6T 3.9T 4.8T 46% /tank
zroot/home 185G 72G 113G 40% /home
===== PERSISTENT MOUNTS (/etc/fstab, non-comment) =====
UUID="3D9B-8E0C" /boot/efi vfat defaults 0 0
===== TARGETED DATA PATHS =====
/tank (total: 3.9T)
total 20
drwxrwxrwx 7 root root 7 2026-06-14 16:11 .
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
drwxrwxr-x 31 llmuser llm 38 2026-07-14 09:27 aimodels
drwxrwxr-x 3 llmuser llm 3 2025-09-08 13:21 comfy
drwxrwxr-x 4 llmuser llmuser 4 2026-04-11 23:17 kokoro
drwxr-xr-x 5 r18clip r18clip 6 2026-06-14 16:20 r18-clip-caption
drwxrwxr-x 5 lkraven lkraven 6 2025-09-10 18:35 vibevoice
/opt (total: 16G)
total 79
drwxrwxrwx 13 root root 13 2026-04-17 23:00 .
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
drwx--x--x 4 root root 4 2025-09-02 12:53 containerd
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 docker
drwxrwxr-x 8 llmuser llm 17 2026-04-17 17:50 heretic
drwxrwxr-x 15 llmuser llmuser 36 2026-04-11 23:24 Kokoro-FastAPI
drwxrwxr-x 20 llmuser llmuser 39 2025-10-09 21:42 LibreChat
drwxrwxr-x 26 llmuser llmuser 56 2025-09-04 21:01 llama.cpp
drwxrwxr-x 13 llmuser llmuser 25 2025-09-02 12:50 llama-swap
drwxrwxr-x 3 llmuser llmuser 11 2026-04-18 15:48 llmcompressor
drwxr-xr-x 4 root root 4 2025-08-30 22:54 nvidia
drwxrwxr-x 6 llmuser llmuser 15 2025-09-24 09:53 parakeet-tdt-0.6b-v2-fastapi
drwxrwxr-x 2 llmuser llmuser 6 2025-09-05 10:03 uv
/opt/docker (total: 110M)
total 18
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 .
drwxrwxrwx 13 root root 13 2026-04-17 23:00 ..
drwxrwxr-x 24 lkraven lkraven 24 2026-07-14 14:50 compose
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 conf
/opt/docker/compose (total: 110M)
total 68
drwxrwxr-x 24 lkraven lkraven 24 2026-07-14 14:50 .
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 18:47 beszel-agent-ana
drwxrwxr-x 2 lkraven lkraven 8 2026-07-16 08:59 char-rp-gguf
drwxr-xr-x 2 lkraven lkraven 4 2025-09-05 20:36 comfyui
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:47 dockge
drwxr-xr-x 2 lkraven lkraven 4 2026-04-20 20:44 dozzle-agent-ana
drwxrwxr-x 3 lkraven lkraven 6 2026-07-16 09:13 heretic2-charrp-reasoning
drwxrwxr-x 3 lkraven lkraven 4 2026-04-11 23:17 kokoro
drwxr-xr-x 2 lkraven lkraven 6 2026-06-12 15:17 llama-swap
drwxr-xr-x 2 lkraven lkraven 6 2026-06-18 23:40 mistral-medium-3.5
drwxr-xr-x 2 lkraven lkraven 5 2026-06-15 18:40 mistral-small-4
drwxr-xr-x 2 root root 4 2026-06-17 22:30 mistral-small-4-heretic
drwxr-xr-x 2 root root 4 2026-07-08 01:10 ms32-24b-angel
drwxr-xr-x 2 lkraven lkraven 4 2025-09-24 12:32 parakeet
drwxr-xr-x 2 lkraven lkraven 5 2026-06-19 01:52 qwen3.5-122b
drwxrwxr-x 2 lkraven lkraven 4 2026-06-13 08:34 qwen35-vl
drwxrwxr-x 2 lkraven lkraven 8 2026-07-16 09:17 qwen36-27b-aeon
drwxr-xr-x 2 lkraven lkraven 5 2026-06-15 20:30 qwen36-vl
drwxr-xr-x 2 lkraven lkraven 5 2026-07-14 16:43 qwen-image-bench
drwxr-xr-x 2 lkraven lkraven 6 2026-07-14 19:56 qwopus3.5-122b
drwxr-xr-x 2 lkraven lkraven 5 2026-07-14 16:43 selene
drwxr-xr-x 2 lkraven lkraven 4 2025-09-10 18:03 vibevoice
drwxr-xr-x 2 lkraven lkraven 13 2026-07-16 10:45 vllm
/opt/docker/conf (total: 14K)
total 2
drwxrwxr-x 4 lkraven lkraven 4 2026-06-04 00:26 .
drwxrwxr-x 4 lkraven lkraven 4 2025-09-03 20:19 ..
drwxr-xr-x 2 lkraven lkraven 3 2026-06-05 00:22 llama-swap
drwxr-xr-x 2 lkraven lkraven 2 2026-06-04 00:37 vllm
/var/lib/docker (total: 8.5K)
/srv (total: 1.0K)
total 10
drwxr-xr-x 2 root root 3 2026-06-14 16:11 .
drwxr-xr-x 18 root root 26 2026-03-29 16:08 ..
lrwxrwxrwx 1 root root 22 2026-06-14 16:11 r18-clip-caption -> /tank/r18-clip-caption
===== DOCKER =====
Server: 29.3.1 Client: 29.3.1
----- docker info -----
Containers: 24 (running 11, paused 0, stopped 13)
Images: 53
Runtimes: map[io.containerd.runc.v2:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} nvidia:{{nvidia-container-runtime [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]} runc:{{runc [] map[]} map[org.opencontainers.runtime-spec.features:{"ociVersionMin":"1.0.0","ociVersionMax":"1.2.1","hooks":["prestart","createRuntime","createContainer","startContainer","poststart","poststop"],"mountOptions":["async","atime","bind","defaults","dev","diratime","dirsync","exec","iversion","lazytime","loud","mand","noatime","nodev","nodiratime","noexec","noiversion","nolazytime","nomand","norelatime","nostrictatime","nosuid","nosymfollow","private","ratime","rbind","rdev","rdiratime","relatime","remount","rexec","rnoatime","rnodev","rnodiratime","rnoexec","rnorelatime","rnostrictatime","rnosuid","rnosymfollow","ro","rprivate","rrelatime","rro","rrw","rshared","rslave","rstrictatime","rsuid","rsymfollow","runbindable","rw","shared","silent","slave","strictatime","suid","symfollow","sync","tmpcopyup","unbindable"],"linux":{"namespaces":["cgroup","ipc","mount","network","pid","time","user","uts"],"capabilities":["CAP_CHOWN","CAP_DAC_OVERRIDE","CAP_DAC_READ_SEARCH","CAP_FOWNER","CAP_FSETID","CAP_KILL","CAP_SETGID","CAP_SETUID","CAP_SETPCAP","CAP_LINUX_IMMUTABLE","CAP_NET_BIND_SERVICE","CAP_NET_BROADCAST","CAP_NET_ADMIN","CAP_NET_RAW","CAP_IPC_LOCK","CAP_IPC_OWNER","CAP_SYS_MODULE","CAP_SYS_RAWIO","CAP_SYS_CHROOT","CAP_SYS_PTRACE","CAP_SYS_PACCT","CAP_SYS_ADMIN","CAP_SYS_BOOT","CAP_SYS_NICE","CAP_SYS_RESOURCE","CAP_SYS_TIME","CAP_SYS_TTY_CONFIG","CAP_MKNOD","CAP_LEASE","CAP_AUDIT_WRITE","CAP_AUDIT_CONTROL","CAP_SETFCAP","CAP_MAC_OVERRIDE","CAP_MAC_ADMIN","CAP_SYSLOG","CAP_WAKE_ALARM","CAP_BLOCK_SUSPEND","CAP_AUDIT_READ","CAP_PERFMON","CAP_BPF","CAP_CHECKPOINT_RESTORE"],"cgroup":{"v1":true,"v2":true,"systemd":true,"systemdUser":true,"rdma":true},"seccomp":{"enabled":true,"actions":["SCMP_ACT_ALLOW","SCMP_ACT_ERRNO","SCMP_ACT_KILL","SCMP_ACT_KILL_PROCESS","SCMP_ACT_KILL_THREAD","SCMP_ACT_LOG","SCMP_ACT_NOTIFY","SCMP_ACT_TRACE","SCMP_ACT_TRAP"],"operators":["SCMP_CMP_EQ","SCMP_CMP_GE","SCMP_CMP_GT","SCMP_CMP_LE","SCMP_CMP_LT","SCMP_CMP_MASKED_EQ","SCMP_CMP_NE"],"archs":["SCMP_ARCH_AARCH64","SCMP_ARCH_ARM","SCMP_ARCH_MIPS","SCMP_ARCH_MIPS64","SCMP_ARCH_MIPS64N32","SCMP_ARCH_MIPSEL","SCMP_ARCH_MIPSEL64","SCMP_ARCH_MIPSEL64N32","SCMP_ARCH_PPC","SCMP_ARCH_PPC64","SCMP_ARCH_PPC64LE","SCMP_ARCH_RISCV64","SCMP_ARCH_S390","SCMP_ARCH_S390X","SCMP_ARCH_X32","SCMP_ARCH_X86","SCMP_ARCH_X86_64"],"knownFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"],"supportedFlags":["SECCOMP_FILTER_FLAG_TSYNC","SECCOMP_FILTER_FLAG_SPEC_ALLOW","SECCOMP_FILTER_FLAG_LOG"]},"apparmor":{"enabled":true},"selinux":{"enabled":true},"intelRdt":{"enabled":true},"mountExtensions":{"idmap":{"enabled":true}}},"annotations":{"io.github.seccomp.libseccomp.version":"2.6.0","org.opencontainers.runc.checkpoint.enabled":"true","org.opencontainers.runc.commit":"v1.3.4-0-gd6d73eb8","org.opencontainers.runc.version":"1.3.4\n"},"potentiallyUnsafeConfigAnnotations":["bundle","org.systemd.property.","org.criu.config"]}]}]
Default runtime: runc
Storage driver: overlay2
Root dir: /var/lib/docker
Server version: 29.3.1
----- running containers -----
NAMES IMAGE STATUS PORTS
vllm-granite vllm/vllm-openai:latest Up 6 days (healthy) 0.0.0.0:8004->8000/tcp, [::]:8004->8000/tcp
vllm-aeon-gen vllm/vllm-openai:latest Up 6 days (healthy) 0.0.0.0:8015->8000/tcp, [::]:8015->8000/tcp
vllm-charrp-reasoning-nvfp4 vllm/vllm-openai:v0.24.0 Up 6 days (healthy) 0.0.0.0:8018->8000/tcp, [::]:8018->8000/tcp
llama-charrp ghcr.io/mostlygeek/llama-swap:cuda Up 6 days (healthy) 0.0.0.0:8016->8080/tcp, [::]:8016->8080/tcp
vllm-reward vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8003->8000/tcp, [::]:8003->8000/tcp
vllm-embed vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8001->8000/tcp, [::]:8001->8000/tcp
vllm-rerank vllm/vllm-openai:latest Up 7 days (healthy) 0.0.0.0:8002->8000/tcp, [::]:8002->8000/tcp
vllm-selene vllm/vllm-openai Up 7 days (healthy) 0.0.0.0:8011->8000/tcp, [::]:8011->8000/tcp
dockge louislam/dockge:latest Up 5 weeks (healthy) 0.0.0.0:5001->5001/tcp, [::]:5001->5001/tcp
dozzle-agent amir20/dozzle:latest Up 5 weeks 0.0.0.0:7007->7007/tcp, 8080/tcp
beszel-agent henrygd/beszel-agent:latest Up 5 weeks (healthy)
----- all containers -----
NAMES IMAGE STATUS
vllm-granite vllm/vllm-openai:latest Up 6 days (healthy)
vllm-aeon-gen vllm/vllm-openai:latest Up 6 days (healthy)
vllm-charrp-reasoning-nvfp4 vllm/vllm-openai:v0.24.0 Up 6 days (healthy)
llama-charrp ghcr.io/mostlygeek/llama-swap:cuda Up 6 days (healthy)
vllm-qwopus35-122b vllm/vllm-openai:latest Created
vllm-aeon-rp vllm/vllm-openai:latest Created
llama-charrp-reasoning llamacpp-charrp:custom-latest Created
vllm-qwen-image-bench vllm/vllm-openai:latest Exited (0) 6 days ago
vllm-reward vllm/vllm-openai:latest Up 7 days (healthy)
vllm-embed vllm/vllm-openai:latest Up 7 days (healthy)
vllm-rerank vllm/vllm-openai:latest Up 7 days (healthy)
vllm-selene vllm/vllm-openai Up 7 days (healthy)
vllm-heretic2-modelopt-quant vllm/vllm-openai:v0.24.0 Exited (0) 8 days ago
vllm-heretic2-cg-quant vllm/vllm-openai:v0.24.0 Exited (0) 8 days ago
rp-dl6 vllm/vllm-openai:latest Exited (0) 2 weeks ago
rp-dl5 vllm/vllm-openai:latest Exited (0) 2 weeks ago
aeon-t1-sft aeon-trainer:latest Exited (0) 2 weeks ago
vllm-deckard-40b vllm/vllm-openai:v0.23.0 Exited (0) 3 weeks ago
selene-dl vllm/vllm-openai:v0.22.0 Exited (0) 5 weeks ago
mistral-dl f37691f675bb Exited (0) 5 weeks ago
r18-staging-dl df7be4c4d818 Exited (0) 5 weeks ago
dockge louislam/dockge:latest Up 5 weeks (healthy)
dozzle-agent amir20/dozzle:latest Up 5 weeks
beszel-agent henrygd/beszel-agent:latest Up 5 weeks (healthy)
----- networks -----
NAME DRIVER SCOPE
bridge bridge local
host host local
kokoro-tts-gpu_default bridge local
librechat_default bridge local
llama-swap_default bridge local
none null local
traefik-net bridge local
----- networks (external, non-default — worth knowing for compose external: true) -----
kokoro-tts-gpu_default
librechat_default
llama-swap_default
traefik-net
----- named volumes -----
VOLUME NAME DRIVER
beszel-agent-ana_beszel_agent_data local
dockge_dockge_data local
dozzle-agent-ana_dozzle_agent_data local
parakeet_parakeet_cache local
searxng_searxng-data local
----- compose projects currently running -----
beszel-agent-ana
char-rp-gguf
dockge
dozzle-agent-ana
heretic2-charrp-reasoning
qwen36-27b-aeon
selene
vllm
===== COMPOSE FILES (/opt/docker/compose/) =====
>>> /opt/docker/compose/beszel-agent-ana/compose.yaml
# Beszel — lightweight server/container monitoring.
#
# Hub: single web UI with the SQLite store. Agents: per-host metric collectors
# that the hub pulls from over SSH.
#
# Multi-host layout via compose profiles:
# COMPOSE_PROFILES=hub → hub only (ana-docker)
# COMPOSE_PROFILES=hub,agent → hub + local agent on the same host
# COMPOSE_PROFILES=agent → agent only (ana-ml2, nh3-docker,
# esh-docker-vm, vm-esh-nas)
#
# The agent uses network_mode: host so it sees real host CPU/mem/net/disk
# counters rather than container-scoped ones — that's why it can't share
# the tnet network with the hub.
#
# All tunables live in .env — edit that, not this file.
services:
beszel:
image: henrygd/beszel:${BESZEL_VERSION}
container_name: beszel
profiles: [hub]
restart: unless-stopped
ports:
- "${BESZEL_PORT}:8090"
volumes:
- beszel_data:/beszel_data
healthcheck:
# Hub image is distroless — no wget/curl. Use the bundled `/beszel`
# binary's built-in health subcommand (https://beszel.dev/guide/healthchecks).
test: ["CMD", "/beszel", "health", "--url", "http://localhost:8090"]
interval: 120s
timeout: 10s
retries: 3
start_period: 15s
networks:
- tnet
labels:
- homepage.group=Monitoring
- homepage.name=Beszel
- homepage.icon=mdi-chart-line
- homepage.description=Server + container monitoring
- homepage.href=http://10.250.50.70:${BESZEL_PORT}
beszel-agent:
image: henrygd/beszel-agent:${BESZEL_VERSION}
container_name: beszel-agent
profiles: [agent]
restart: unless-stopped
network_mode: host
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- beszel_agent_data:/var/lib/beszel-agent
environment:
# Agent auth has two modes (v0.13+ supports both side-by-side):
# - KEY-mode: agent listens, hub connects inbound over SSH using KEY.
# Requires BESZEL_HUB_KEY in .env.
# - Token-mode: agent initiates an outbound connection to HUB_URL
# using TOKEN. Easier through NAT. Requires HUB_URL + BESZEL_TOKEN.
# Leave unused ones empty ("") in .env; both can be set simultaneously.
- PORT=${BESZEL_AGENT_PORT:-45876}
- KEY=${BESZEL_HUB_KEY:-}
- HUB_URL=${HUB_URL:-}
- TOKEN=${BESZEL_TOKEN:-}
- EXTRA_FILESYSTEMS=${BESZEL_EXTRA_FS:-}
healthcheck:
# Agent image ships the `/agent` binary with a `health` subcommand.
# Verifies the agent process is up — not that the hub can reach it.
test: ["CMD", "/agent", "health"]
interval: 120s
timeout: 10s
retries: 3
start_period: 15s
volumes:
beszel_data:
beszel_agent_data:
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/char-rp-gguf/compose.yaml
# char-rp-gguf — dedicated GGUF character-RP seat on ana-ml2 GPU 0, REPLACING the
# broken ms32-24b-angel NVFP4 serve (garbage output — bad self-quant W4A4).
#
# Two co-located llama.cpp (llama-server) instances on GPU 0, served alongside the
# 35B-A3B heretic `gen` (qwen36-27b-aeon stack, :8015):
#
# llama-charrp (:8016, gateway char-rp) — TheDrummer Magidonia-24B-v4.3 Q6_K.
# Magistral (Mistral) dark-romantasy RP tune. NON-thinking PROSE seat: elite
# literary prose, zero refusal, ~65 tok/s, precise POV/instruction adherence.
#
# llama-charrp-reasoning (:8018, gateway char-rp-reasoning) — ArliAI QwQ-32B-RpR-v4 Q5_K_M.
# QwQ reasoning RP tune whose reasoning DATA was generated with QwQ-ABLITERATED
# → it does NOT re-censor in the think phase (the exact failure mode that killed
# the Pantheon/DeepSeek-distilled reasoners: they reason themselves into refusals
# inside <think>). llama.cpp MANAGES QwQ reasoning natively: --reasoning on
# surfaces the trace in reasoning_content (clean prose in content, no <think>
# leak), --reasoning-budget caps the chain-of-thought. ~50 tok/s @ Q5_K_M.
#
# WHY GGUF/llama.cpp (not vLLM NVFP4): sidesteps BOTH traps that killed the Angel serve
# — the vLLM NVFP4 self-quant breakage AND the Mistral-tokenizer/vision crash. llama.cpp
# handles Mistral + QwQ tokenizers natively. NEVER Ollama (banned fleet-wide).
#
# WHY TWO models (not one): no single dense 24-32B is BOTH an elite non-thinking prose
# seat AND a clean managed-reasoning seat on llama.cpp. Magidonia's Magistral [THINK]
# discipline is loose (won't reliably close [/THINK] on substantive reasoning → prose
# bleeds into reasoning_content, content empties); Cydonia-R1's <think> is emergent, so
# llama.cpp can't manage/cap it → runaway CoT that never reaches prose. QwQ's template
# opens <think> natively → llama.cpp manages+caps it. So: best-of-breed per seat.
# ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning): point both services
# at Magidonia via CHARRP_REASONING_MODEL in .env and blank CHARRP_REASONING_EXTRA_*.
#
# ALTERNATE prose model: PaintedFantasy-v4.1-24B (also Magistral, more literary flair
# but looser POV adherence) — set CHARRP_MODEL in .env. All candidate GGUFs are
# pre-pulled to /tank/aimodels/llm/rp/.
#
# VRAM (GPU 0, co-resident with gen ~38G): Magidonia Q6 ~19G + RpR-v4 Q5 ~23G + KV/
# compute ~6-8G = ~85-88G / 97G (~9-12G margin). Keep ctx modest; drop CHARRP_*_CTX
# to 8192 in .env if warmup bites. depends_on sequences char-rp first.
#
# API auth: blank (LAN-internal on the GPU host; matches API_KEY= in the AEON stack /
# gateway VLLM_API_KEY). llama-server ignores the gateway's api_key when none is set.
#
# All tunables live in .env — edit that, not this file.
name: char-rp-gguf
services:
# ── PROSE seat — non-thinking. gateway char-rp. ──
llama-charrp:
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
container_name: ${CHARRP_CONTAINER:-llama-charrp}
restart: unless-stopped
runtime: nvidia
ports:
- "${CHARRP_PORT:-8016}:8080"
volumes:
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
environment:
# Pin to GPU 0 (the on-demand large-model card; the always-on vLLM trio owns GPU 1).
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
entrypoint: ["/app/llama-server"]
command:
- --model
- /models/${CHARRP_MODEL:-rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf}
- --host
- 0.0.0.0
- --port
- "8080"
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_CTX:-98304}"
- --flash-attn
- on
# q8_0 KV cache ~halves KV VRAM (8-bit, near-lossless) → ~2x the context per GB.
# Mistral/Magistral handles q8 KV cleanly. Set f16 in .env to disable.
- --cache-type-k
- ${CHARRP_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_KV_TYPE:-q8_0}
- --jinja
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 240s
networks:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=char-rp (Magidonia-24B GGUF)
- homepage.icon=mdi-drama-masks
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
# ── REASONING seat — NEO-CODE = Heretic2-Thinking (Qwen3.6-27B) managed thinking. gateway char-rp-reasoning. ──
llama-charrp-reasoning:
# ⚠️ CUSTOM llama.cpp build (master 6eddde0 + unmerged PR #25544). Needed for two reasons:
# (1) recent master parses Qwen3.6's native qwen3_coder tool-call format (stock b8840 predates it —
# the <tool_call><function=..><parameter=..> XML is Qwen3.5/3.6-native, NOT an OpenHands quirk);
# (2) PR #25544 multi-terminator reasoning-budget fix (Worldtree #355) — belt-and-suspenders now that
# NEO-CODE shows 0.0 runaway (R36 gate), but keep it. DO NOT revert to stock until #25544 merges.
# Build recipe + why + rollback: ./llamacpp-custom/README.md.
# Rollback: set LLAMA_REASONING_IMAGE=ghcr.io/mostlygeek/llama-swap:cuda in .env + recreate.
image: ${LLAMA_REASONING_IMAGE:-llamacpp-charrp:custom-latest}
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
restart: unless-stopped
runtime: nvidia
# Sequence AFTER the prose seat is healthy so the two GPU-0 allocations don't race.
depends_on:
llama-charrp:
condition: service_healthy
ports:
- "${CHARRP_REASONING_PORT:-8018}:8080"
volumes:
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
entrypoint: ["/app/llama-server"]
command:
- --model
- /models/${CHARRP_REASONING_MODEL:-rp/Qwen3.6-27B-NEO-CODE-HERE-2T-OT-Q5_K_M.gguf}
- --host
- 0.0.0.0
- --port
- "8080"
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_REASONING_CTX:-40960}"
- --flash-attn
- on
# NEO-CODE = Qwen3.6-27B GDN-hybrid (16 of 64 layers cache KV → KV cheap); native ctx 262144
# (256K). Full 256K @ q8_0 KV ≈ 8.6G, fits GPU0 w/ ~3.8G margin. q8_0 coherent; f16 in .env if gibberish.
- --cache-type-k
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --jinja
# NEO-CODE's Qwen3.6 template natively opens <think> → llama.cpp manages the reasoning
# (trace to reasoning_content, content stays clean prose); --reasoning-budget caps the CoT.
# (R36 gate 2026-07-14: NEO-CODE composite 0.922 tool-calling + 0.0 runaway — beat Deckard
# 0.08/0.80 and gen-reasoning 0.856. Budget held at 400: latency-coupled to soong's client timeout.)
- --reasoning
- on
- --reasoning-format
- deepseek
- --reasoning-budget
- "${CHARRP_REASONING_BUDGET:-400}"
# Sampler defaults per the DavidAU/Qwen3.6 model card (thinking-mode, general tasks): temp 1.0,
# top_p 0.95, top_k 20, min_p 0.0, no rep-penalty, no DRY (DRY was a QwQ/Deckard looping band-aid
# NEO-CODE doesn't need). All tunable via .env. NOTE: 0.922 tool-gate was on the OLD Deckard
# samplers (effective temp~0.8 + DRY); re-validate tools + slop on these card samplers.
- --temp
- "${CHARRP_REASONING_TEMP:-1.0}"
- --top-p
- "${CHARRP_REASONING_TOP_P:-0.95}"
- --top-k
- "${CHARRP_REASONING_TOP_K:-20}"
- --min-p
- "${CHARRP_REASONING_MIN_P:-0.0}"
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI - Dormant
- homepage.name=char-rp-reasoning (QwQ-32B RpR-v4 GGUF)
- homepage.icon=mdi-brain
- homepage.description=Dark-romantasy RP reasoning seat, managed CoT (llama.cpp, ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${CHARRP_REASONING_PORT:-8018}
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/comfyui/compose.yaml
services:
comfyui:
runtime: nvidia
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities:
- gpu
- compute
- utility
ports:
- 8188:8188
image: mmartial/comfyui-nvidia-docker:ubuntu24_cuda13.0-latest
networks:
- tnet
volumes:
- /tank/comfy/run:/comfy/mnt
- /tank/aimodels/img/comfy:/basedir
#user: 1001:1002
environment:
- WANTED_UID=1001
- WANTED_GID=1002
- BASE_DIRECTORY=/basedir
- SECURITY_LEVEL=weak
- NVIDIA_VISIBLE_DEVICES=all
- NVIDIA_DRIVER_CAPABILITIES=all
labels:
- homepage.group=AI Systems
- homepage.name=ComfyUI
- homepage.icon=mdi-panorama-variant-outline
- homepage.description=ComfyUI Image Gen (ana-ml2)
- homepage.href=http://10.250.50.54:8188
restart: unless-stopped
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/dockge/compose.yaml
# Dockge — per-host Docker Compose UI (https://dockge.kuma.pet/).
#
# One instance runs on every Docker host so the compose dir is manageable
# from a browser. Each host sets DOCKGE_HOST_LABEL + DOCKGE_HOST_IP in its
# .env so the homepage card points at the right place.
#
# All tunables live in .env — edit that, not this file.
services:
dockge:
image: louislam/dockge:${DOCKGE_VERSION:-latest}
container_name: dockge
restart: unless-stopped
ports:
- "${DOCKGE_PORT:-5001}:5001"
volumes:
- /var/run/docker.sock:/var/run/docker.sock
- dockge_data:/app/data
- /opt/docker:/opt/docker
environment:
- DOCKGE_STACKS_DIR=/opt/docker/compose
networks:
- tnet
labels:
- homepage.group=Service Networking
- homepage.name=Dockge (${DOCKGE_HOST_LABEL})
- homepage.icon=sh-dockge.png
- homepage.description=Compose UI on ${DOCKGE_HOST_LABEL}
- homepage.href=http://${DOCKGE_HOST_IP}:${DOCKGE_PORT:-5001}
volumes:
dockge_data:
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/dozzle-agent-ana/compose.yaml
# Dozzle — container log viewer.
#
# Multi-host layout via compose profiles:
# COMPOSE_PROFILES=hub → runs the web UI (deploy on ana-docker)
# COMPOSE_PROFILES=agent → runs the remote agent (deploy on ana-ml2)
#
# Same compose.yaml on both servers; per-host `.env` picks the profile.
#
# All tunables live in .env — edit that, not this file.
services:
dozzle:
image: amir20/dozzle:${DOZZLE_VERSION}
container_name: dozzle
profiles: [hub]
restart: unless-stopped
ports:
- "${DOZZLE_PORT}:8080"
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- dozzle_data:/data
environment:
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
- DOZZLE_REMOTE_AGENT=${DOZZLE_REMOTE_AGENT:-}
- DOZZLE_AUTH_PROVIDER=${DOZZLE_AUTH_PROVIDER:-none}
- DOZZLE_USERNAME=${DOZZLE_USERNAME:-}
- DOZZLE_PASSWORD=${DOZZLE_PASSWORD:-}
healthcheck:
test: ["CMD", "/dozzle", "healthcheck"]
interval: 30s
timeout: 10s
retries: 3
start_period: 15s
networks:
- tnet
labels:
- homepage.group=Monitoring
- homepage.name=Dozzle
- homepage.icon=mdi-text-box-search
- homepage.description=Container logs (ana-docker + ana-ml2)
- homepage.href=http://10.250.50.70:${DOZZLE_PORT}
dozzle-agent:
image: amir20/dozzle:${DOZZLE_VERSION}
container_name: dozzle-agent
profiles: [agent]
restart: unless-stopped
command: agent
ports:
- "${DOZZLE_AGENT_BIND:-0.0.0.0}:${DOZZLE_AGENT_PORT}:7007"
volumes:
- /var/run/docker.sock:/var/run/docker.sock:ro
- dozzle_agent_data:/data
environment:
- DOZZLE_HOSTNAME=${DOZZLE_HOSTNAME}
networks:
- tnet
volumes:
dozzle_data:
dozzle_agent_data:
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/heretic2-charrp-reasoning/compose.yaml
# heretic2-charrp-reasoning — modelopt NVFP4 + native MTP fast char-rp-reasoning seat on
# ana-ml2 GPU0, replacing the GGUF NEO-CODE reasoning seat (llama-charrp-reasoning, now retired).
# Same Heretic2/NEO-CODE model; ~77 tok/s (~1.3x over GGUF) via qwen3_5_mtp spec-decode.
#
# ⚠️ REQUIRES the MTP workaround: vLLM 0.24.0 doesn't propagate modelopt exclude_modules to the
# spec-decode DRAFT model, so the BF16 mtp head gets quantized -> shape crash. The mounted
# sitecustomize.py (conf/mtp-workaround/) force-skips mtp.* in is_layer_skipped. Without it the
# engine dies at load. Full recipe: eshpfi docs/runbooks/heretic2-nvfp4-mtp-seat.md.
#
# Co-located on GPU0 with vllm-aeon-gen (gen) + llama-charrp (char-rp). VRAM: NVFP4 27B weights
# ~26GB + KV. util 0.30 fits the ~33GB free alongside gen+char-rp -> max-model-len capped at
# 32768 (the GGUF seat did 256K on lighter Q5 weights; NVFP4 is heavier, so context is reduced
# until VRAM is rebalanced). Tunables in .env.
name: heretic2-charrp-reasoning
services:
vllm-charrp-reasoning:
image: ${REASONING_IMAGE:-vllm/vllm-openai:v0.24.0}
container_name: ${REASONING_CONTAINER:-vllm-charrp-reasoning-nvfp4}
restart: unless-stopped
ipc: host
ports:
- "${REASONING_PORT:-8018}:8000"
volumes:
- /tank/aimodels:/tank/aimodels
# The MTP draft-model quant workaround (sitecustomize.py). PYTHONPATH loads it in the
# engine-core subprocess. See runbook landmine #4.
- ./conf/mtp-workaround:/mtp-workaround:ro
environment:
- PYTHONPATH=/mtp-workaround
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
- VLLM_API_KEY=${API_KEY:-}
command:
- ${REASONING_MODEL:-/tank/aimodels/heretic2-nvfp4-work/heretic2-modelopt-nvfp4-mtp}
- --quantization
- modelopt
- --speculative-config
- '{"method": "qwen3_5_mtp", "num_speculative_tokens": ${SPEC_TOKENS:-3}}'
- --language-model-only
- --mamba-cache-dtype
- float32
- --reasoning-parser
- qwen3
- --tool-call-parser
- qwen3_coder
- --enable-auto-tool-choice
- --served-model-name
- char-rp-reasoning
- --max-model-len
- "${REASONING_MAX_MODEL_LEN:-32768}"
- --max-num-seqs
- "${REASONING_MAX_NUM_SEQS:-4}"
- --gpu-memory-utilization
- "${REASONING_GPU_MEM_UTIL:-0.30}"
- --kv-cache-dtype
- fp8
- --trust-remote-code
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${REASONING_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 600s
networks:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=char-rp-reasoning (Heretic2 NVFP4+MTP)
- homepage.icon=mdi-rocket-launch
- homepage.description=NEO-CODE Heretic2 NVFP4 + native MTP, ~77 tok/s (ana-ml2 GPU0)
- homepage.href=http://10.250.50.54:${REASONING_PORT:-8018}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/kokoro/compose.yaml
name: kokoro-tts
services:
kokoro-tts:
container_name: kokoro-tts
build:
context: ./Kokoro-FastAPI
dockerfile: docker/gpu/Dockerfile
volumes:
- /tank/kokoro/models:/app/api/src/models
- /tank/kokoro/output:/app/output
ports:
- "8765:8880"
environment:
- PYTHONPATH=/app:/app/api
- USE_GPU=true
- PYTHONUNBUFFERED=1
- DOWNLOAD_MODEL=false
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Kokoro TTS
- homepage.icon=mdi-waveform
- homepage.description=Kokoro FastAPI TTS (OpenAI-compatible)
- homepage.href=http://10.250.50.54:8765
restart: unless-stopped
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/llama-swap/compose.yaml
# llama-swap — GGUF model server with on-demand model swapping.
#
# Proxies OpenAI-compatible API requests to llama.cpp server instances
# and swaps which model is loaded into VRAM per request. Runs on
# ana-ml2 using both GPUs dynamically (no explicit device pinning —
# llama-swap picks per-model-definition).
#
# Model definitions live in /opt/docker/conf/llama-swap/config.yaml on
# the server. Canonical copy of that config is config.yaml in this
# workspace; deploy with scp + `docker compose restart` or the script
# at the bottom of README.md.
#
# All tunables live in .env — edit that, not this file.
services:
llama-swap:
image: ghcr.io/mostlygeek/llama-swap:${LLAMA_SWAP_VERSION}
container_name: llama-swap
restart: unless-stopped
stdin_open: true
tty: true
runtime: nvidia
ports:
- "${LLAMA_SWAP_PORT}:8080"
volumes:
- /opt/docker/conf/llama-swap/config.yaml:/app/config.yaml
- ${MODELS_DIR}:/models
- ${HF_CACHE_DIR}:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
# Pin to GPU 0 — the reserved card for on-demand large-model hot-loads.
# The always-on vLLM services (granite + embed/rerank/reward) own GPU 1;
# keeping llama-swap off GPU 1 stops a hot-loaded model from contending
# with them. llama.cpp then sees only GPU 0 (cuda:0), so --n-gpu-layers
# 999 loads there with no per-model device targeting needed.
- NVIDIA_VISIBLE_DEVICES=${LLAMA_SWAP_GPU:-0}
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/ >/dev/null || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 30s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=llama-swap
- homepage.icon=mdi-swap-horizontal
- homepage.description=GGUF model swapper (llama.cpp; ana-ml2)
- homepage.href=http://10.250.50.54:${LLAMA_SWAP_PORT}
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/mistral-medium-3.5/compose.yaml
# mistral-medium-3.5 — RecViking/Mistral-Medium-3.5-128B-NVFP4 on ana-ml2 GPU 0
# as a TEMPORARY speed-check tenant, DISPLACING mistral-small-4 (operator 2026-06-19:
# "pull the NVFP4 model and serve it ... displace mistral-small-4 for now ... I want
# to check it").
#
# Model: Mistral Medium 3.5, 128B `Mistral3ForConditionalGeneration` (mistral3,
# multimodal), RecViking's NVFP4 (compressed-tensors / nvfp4-pack-quantized), HF format.
#
# SERVING (per RecViking's model card): vLLM NIGHTLY loads the HF-format NVFP4 weights
# DIRECTLY — no Mistral native-convert (unlike Small 4) — via the FlashInfer Cutlass
# NVFP4 kernel + TURBOQUANT 4-bit KV. RecViking used TP=4; the ~70 GB NVFP4 fits one
# 96 GB Blackwell card, so we run TP=1 on GPU 0. Context trimmed to 32K (speed check,
# not full 256K) so KV fits comfortably on one card.
#
# DISPLACEMENT: GPU 0 holds only one mistral-class model. Bring this up only after
# downing the live mistral-small-4-heretic stack. REVERT = `docker compose down` this,
# then `up -d` /opt/docker/compose/mistral-small-4-heretic (restores the Worldtree
# character backend). Served under its OWN name (mistral-medium-3.5), NOT
# mistral-small-4 — no stale-alias (the Worldtree character `mistral-small-4` route
# 404s while this is up; that's the "for now").
#
# EAGLE: RecViking's repo has no EAGLE head. The official native head
# (mistralai/Mistral-Medium-3.5-128B-EAGLE) is staged at /tank/aimodels/
# mistral-medium-3.5-eagle, but wiring spec-decode (native head + HF base, mistral3
# arch, nightly) is a follow-on — off by default. See README.
name: mistral-medium-3.5
services:
vllm-medium35:
image: ${MEDIUM35_IMAGE:-vllm/vllm-openai:nightly}
container_name: ${MEDIUM35_CONTAINER_NAME:-vllm-medium35}
restart: unless-stopped
ipc: host
ports:
- "${MEDIUM35_PORT:-8012}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
# RecViking NVFP4 checkpoint (HF format, read-only).
- /tank/aimodels/mistral-medium-3.5-nvfp4:/model:ro
# Official EAGLE draft head (native FP8, 2-layer) for speculative decoding.
- /tank/aimodels/mistral-medium-3.5-eagle:/eagle:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
command:
- /model
# HF-format load (NO --config-format/--load-format/--tokenizer-mode mistral —
# that's the Small 4 native path; nightly serves this NVFP4 from HF directly).
- --served-model-name
- mistral-medium-3.5
- --host
- 0.0.0.0
- --port
- "8000"
- --tensor-parallel-size
- "1"
- --gpu-memory-utilization
- ${MEDIUM35_GPU_MEM_UTIL:-0.93}
- --max-model-len
- ${MEDIUM35_MAX_MODEL_LEN:-32768}
# TURBOQUANT 4-bit KV (nightly) — RecViking's recommended KV dtype for this NVFP4.
- --kv-cache-dtype
- ${MEDIUM35_KV_CACHE_DTYPE:-turboquant_4bit_nc}
- --max-num-seqs
- ${MEDIUM35_MAX_NUM_SEQS:-16}
- --dtype
- auto
- --enable-prefix-caching
# EAGLE speculative decoding (method eagle, 3 spec tokens per the EAGLE card).
# Draft head mounted at /eagle. Remove these to revert to base-only.
# --enforce-eager: the EAGLE+NVFP4+turboquant path crashes in CUDA-graph replay
# on this nightly; disabling graphs is the workaround (costs some base speed).
- --enforce-eager
- --speculative-config
- '{"model": "/eagle", "num_speculative_tokens": 3, "method": "eagle"}'
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${MEDIUM35_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
# Nightly image + 70 GB NVFP4 load + CUDA/kernel warmup.
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Mistral Medium 3.5 (NVFP4 speed-check)
- homepage.icon=mdi-speedometer
- homepage.description=RecViking Mistral-Medium-3.5-128B NVFP4 on vLLM nightly (ana-ml2 GPU 0, displacing Small 4)
- homepage.href=http://10.250.50.54:${MEDIUM35_PORT:-8012}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/mistral-small-4/compose.yaml
# mistral-small-4 — Mistral-Small-4-119B-2603 (official NVFP4) on ana-ml2 GPU 0.
#
# Mistral Small 4 is a 119B-total / 6.5B-active MoE (128 experts, 4 active),
# 256K context, multimodal, Apache-2.0 (released 2026-03). This serves the
# OFFICIAL NVFP4 checkpoint (mistralai/Mistral-Small-4-119B-2603-NVFP4) — 74.4 GB
# of compressed-tensors (llm-compressor, a vLLM + Red Hat collaboration, day-0
# vLLM support). It is the GPU-0 tenant (the slot formerly reserved for a
# creative-writing pick — operator reassigned 2026-06-15; tune-for-creative-
# writing comes after base-characteristic probing).
#
# WHY NVFP4 (not FP8/bf16): on a SINGLE 96 GB card, NVFP4 (74.4 GB weights) is
# the only variant that fits at TP=1 — FP8 (~119 GB) and bf16 (~238 GB) need both
# GPUs. The card is Blackwell (sm_120) with FP4 tensor cores, so NVFP4 gets a real
# speedup, not just a VRAM save. NOTE: this is the COMPRESSED-TENSORS NVFP4 path
# (vendor-shipped, vLLM-tested) — distinct from the nvidia-ModelOpt NVFP4 MoE
# loader that broke on Qwen3.6 (#44081); different code path, day-0 supported.
#
# WHY TP=1 here: Mistral's official card uses --tensor-parallel-size 2 (their
# reference 80 GB cards can't fit 74.4 GB + context on one). The 96 GB Blackwell
# flips that to single-card: 74.4 GB weights + ~5 GB overhead leaves ~17 GB for
# KV. Mistral Small 4 uses MLA attention (TRITON_MLA) so KV is compressed/cheap —
# big context stays affordable even on a constrained KV pool. max-model-len is
# capped to 131072 on first bring-up (raise toward the native 256K once real KV
# headroom is measured).
#
# vLLM FLOOR: needs >= 0.20 (Mistral Small 4 day-0 support); validated on 0.23.0.
# Do NOT reuse the qwen36-vl 0.19.1 image — it predates this model.
#
# Serve flags mirror Mistral's official command (cited in README), adapted for
# single-card: TP 2->1, util 0.8->0.93, max-len 262144->131072, max-num-seqs
# 128->64. All tunables live in .env — edit that, not this file.
name: mistral-small-4
services:
vllm-mistral4:
image: ${MISTRAL_IMAGE}
container_name: ${MISTRAL_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${MISTRAL_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${MISTRAL_MODEL}
# Pre-quantized NVFP4 (compressed-tensors) — vLLM auto-detects the quant;
# no --quantization flag.
- --served-model-name
- mistral-small-4
- --host
- 0.0.0.0
- --port
- "8000"
- --tensor-parallel-size
- "1"
- --gpu-memory-utilization
- ${MISTRAL_GPU_MEM_UTIL}
- --max-model-len
- ${MISTRAL_MAX_MODEL_LEN}
# MLA attention backend (DeepSeek-style latent KV → compressed, cheap KV).
- --attention-backend
- TRITON_MLA
# Mistral tool-calling + configurable reasoning (per the official card).
- --tool-call-parser
- mistral
- --enable-auto-tool-choice
- --reasoning-parser
- mistral
- --max-num-seqs
- ${MISTRAL_MAX_NUM_SEQS}
# VISION ENABLED. vLLM is pinned to v0.22.0 in .env — the last release BEFORE
# the Mistral multimodal regression (#44911, `MistralCommonImageProcessor has
# no attribute fetch_images`, landed ~0.22.1+; 0.23.0 is affected). v0.22.0
# still has Mistral-Small-4 arch + compressed-tensors NVFP4 support (the
# #44081 ModelOpt-NVFP4 bug on 0.22.0 is a DIFFERENT quant path, doesn't touch
# this compressed-tensors checkpoint). Gives a verified working vision tower
# as the abliteration/tuning baseline. (qwen36 stays on 0.23.0 — separate
# container; it NEEDS 0.23.0 for its ModelOpt NVFP4.)
- --dtype
- auto
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${MISTRAL_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 600s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Mistral Small 4 (NVFP4)
- homepage.icon=mdi-creation
- homepage.description=Mistral-Small-4-119B-2603 MoE (NVFP4) via vLLM (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/mistral-small-4-heretic/compose.yaml
# mistral-small-4-heretic — abliterated Mistral Small 4 (heretic NVFP4) as a
# DROP-IN for the official mistral-small-4 backend.
#
# Serves darkc0de/Mistral-Small-4-119B-2603-heretic, quantized to NVFP4 in-house
# (vision tower kept bf16) and converted to Mistral native format. Built + validated
# 2026-06-17 — see tools/mistral-small4-nvfp4/ for the build pipeline.
#
# WHY a separate stack: GPU0 fits only one mistral-class model (~65-70 GB), so this
# is a backend SWAP, not a co-tenant. Bring it up only after downing the official
# mistral-small-4 stack. It serves under --served-model-name mistral-small-4 on the
# SAME port (8010), so litellm's mistral-small-4 + mistral-small-4-reasoning entries
# route here with NO litellm change. Revert = down this, `up -d` the official stack.
#
# DIFFERENCES vs the official compose (everything else mirrors it for a faithful
# drop-in — TP=1, util 0.93, MLA, reasoning + tool-call parsers, prefix caching):
# - model is a LOCAL native dir (mounted /model), not an HF id, so it needs the
# native loader flags: --config-format/--load-format/--tokenizer-mode mistral.
# - distinct container_name (vllm-mistral4-heretic) so it can be staged without
# colliding with the official container.
name: mistral-small-4-heretic
services:
vllm-mistral4-heretic:
image: ${MISTRAL_IMAGE}
container_name: ${MISTRAL_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${MISTRAL_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
# The in-house heretic native NVFP4 checkpoint (read-only).
- /tank/aimodels/quant-work/heretic-native-nvfp4:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- /model
# Native Mistral format (params.json + consolidated*.safetensors + tekken).
- --config-format
- mistral
- --load-format
- mistral
- --tokenizer-mode
- mistral
# SAME served name as the official → litellm routes here unchanged.
- --served-model-name
- mistral-small-4
- --host
- 0.0.0.0
- --port
- "8000"
- --tensor-parallel-size
- "1"
- --gpu-memory-utilization
- ${MISTRAL_GPU_MEM_UTIL}
- --max-model-len
- ${MISTRAL_MAX_MODEL_LEN}
- --attention-backend
- TRITON_MLA
- --tool-call-parser
- mistral
- --enable-auto-tool-choice
- --reasoning-parser
- mistral
- --max-num-seqs
- ${MISTRAL_MAX_NUM_SEQS}
- --dtype
- auto
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${MISTRAL_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 600s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Mistral Small 4 (heretic NVFP4)
- homepage.icon=mdi-creation
- homepage.description=Abliterated Mistral-Small-4 (heretic NVFP4) drop-in via vLLM (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${MISTRAL_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/ms32-24b-angel/compose.yaml
# ms32-24b-angel — allura-org/MS3.2-24b-Angel (Mistral-Small-3.2-24B RP/fiction finetune) on ana-ml2 GPU 0.
# Serves the char-rp slot (:8016), REPLACING the qwen aeon-rp seat. NVFP4 (compressed-tensors: MLP quantized,
# vision tower + attention + lm_head kept bf16). NON-reasoning RP model — so NO MTP/spec-decode, NO GDN
# mamba-cache, NO reasoning-parser. Mistral tokenizer (--tokenizer-mode mistral, per the model card).
# Served under the aeon-rp names so the LiteLLM gateway char-rp / char-rp-reasoning routing stays transparent.
# Sampling defaults live at the gateway (RP: temp 1.2 / min_p 0.1 / rep 1.05, per the card + community).
# REVERT: `docker compose down` here + `docker compose up -d vllm-aeon-rp` in ../qwen36-27b-aeon.
name: ms32-24b-angel
services:
vllm-angel-rp:
image: ${ANGEL_IMAGE:-vllm/vllm-openai:latest}
container_name: ${ANGEL_CONTAINER_NAME:-vllm-angel-rp}
restart: unless-stopped
ipc: host
ports:
- "${ANGEL_PORT:-8016}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
- ${ANGEL_MODEL:-/tank/aimodels/ms32-24b-angel-nvfp4}:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
command:
- /model
- --served-model-name
- ${ANGEL_SERVED_NAME:-qwen3.6-27b-aeon-rp}
- ${ANGEL_SERVED_NAME_ALT:-qwen3.6-27b-aeon-rp-thinking}
- --host
- 0.0.0.0
- --port
- "8000"
- --quantization
- compressed-tensors
- --tokenizer-mode
- ${ANGEL_TOKENIZER_MODE:-auto}
- --gpu-memory-utilization
- ${ANGEL_GPU_MEM_UTIL:-0.35}
- --max-model-len
- ${ANGEL_MAX_MODEL_LEN:-131072}
- --max-num-seqs
- ${ANGEL_MAX_NUM_SEQS:-4}
- --dtype
- auto
- --kv-cache-dtype
- ${ANGEL_KV_CACHE_DTYPE:-fp8}
- --enable-prefix-caching
- --limit-mm-per-prompt
- '{"image": 0}'
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${ANGEL_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=MS3.2-24B Angel (NVFP4, vision) — char-rp
- homepage.icon=mdi-drama-masks
- homepage.description=allura-org MS3.2-24B Angel, uncensored RP/fiction (char-rp), ana-ml2 GPU 0
- homepage.href=http://10.250.50.54:${ANGEL_PORT:-8016}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/parakeet/compose.yaml
services:
parakeet-stt:
image: parakeet-stt
ports:
- 8300:8000
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities:
- gpu
restart: unless-stopped
volumes:
- parakeet_cache:/root/.cache
networks:
- tnet
env_file:
- .env
labels:
- homepage.group=AI Systems
- homepage.name=Parakeet
- homepage.icon=mdi-talk
- homepage.description=Parakeet STT (ana-ml2)
- homepage.href=http://10.250.50.54:8300
networks:
tnet:
name: traefik-net
external: true
volumes:
parakeet_cache: null
>>> /opt/docker/compose/qwen3.5-122b/compose.yaml
# qwen3.5-122b — bjk110/Qwen3.5-122B-A10B-abliterated-NVFP4 on ana-ml2 GPU 0,
# REPLACING mistral-small-4 (operator 2026-06-19: down the heretic, serve this as
# the new general/`gen` model). Abliterated Qwen3.5 MoE (256 experts, 10B active),
# NVFP4 (compressed-tensors), HF format.
#
# SERVING (per the repo's serving/): Qwen3.5 MoE is a MULTIMODAL arch but this
# checkpoint is text-only, so vLLM needs the repo's text-only PATCH applied before
# startup. We reuse the repo's entrypoint.sh (applies the patch, then runs vLLM) and
# vllm_patches/, mounted from the downloaded model dir — keeps patch+checkpoint
# version-coupled. Thinking split via litellm extra_body chat_template_kwargs
# (enable_thinking) + --reasoning-parser qwen3 (mirrors the qwen3.6-35b-a3b pattern).
#
# DISPLACEMENT: GPU 0 fits one mistral-class model; bring this up only after downing
# mistral-small-4-heretic. REVERT = down this, `up -d` the heretic stack.
#
# NOTE: image is vllm/vllm-openai:latest per the repo (the patch targets latest) —
# MUTABLE tag; pin a digest once a known-good version is established.
#
# Tunables in .env.
name: qwen3.5-122b
services:
vllm-qwen35-122b:
image: ${QWEN35_IMAGE:-vllm/vllm-openai:latest}
container_name: ${QWEN35_CONTAINER_NAME:-vllm-qwen35-122b}
restart: unless-stopped
ipc: host
ports:
- "${QWEN35_PORT:-8013}:8000"
environment:
- ROLE=head
- TP_SIZE=1
- MODEL_CONTAINER_PATH=/models/qwen
- SERVED_MODEL_NAME=${QWEN35_SERVED_NAME:-qwen3.5-122-a10b}
- HOST_PORT=8000
- MAX_MODEL_LEN=${QWEN35_MAX_MODEL_LEN:-131072}
- MAX_NUM_SEQS=${QWEN35_MAX_NUM_SEQS:-8}
- GPU_MEMORY_UTILIZATION=${QWEN35_GPU_MEM_UTIL:-0.95}
- MAX_NUM_BATCHED_TOKENS=${QWEN35_MAX_NUM_BATCHED_TOKENS:-32768}
# --reasoning-parser qwen3 surfaces <think>…</think> as reasoning_content;
# the thinking on/off itself is per-request (litellm chat_template_kwargs).
# --enable-auto-tool-choice + --tool-call-parser: Qwen3.5 emits XML tool calls
# <tool_call><function=NAME><parameter=K>V</parameter></function></tool_call>
# (NOT Hermes JSON), so the parser is qwen3_xml. Without these flags vLLM never
# parses tool calls (tool-calling is broken). The bjk110 repo command omitted them.
- VLLM_EXTRA_ARGS=--reasoning-parser qwen3 --enable-chunked-prefill --enable-auto-tool-choice --tool-call-parser qwen3_xml
- NVIDIA_VISIBLE_DEVICES=0
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
- VLLM_API_KEY=${API_KEY:-}
volumes:
# checkpoint + the repo's entrypoint/patch (downloaded with the model).
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4:/models/qwen:ro
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/serving/entrypoint.sh:/entrypoint.sh:ro
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/vllm_patches:/patches:ro
- /tank/aimodels/qwen3.5-122b-a10b-nvfp4/.cache/vllm:/root/.cache/vllm
entrypoint: ["bash", "/entrypoint.sh"]
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${QWEN35_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Qwen3.5-122B-A10B (abliterated NVFP4)
- homepage.icon=mdi-creation
- homepage.description=Abliterated Qwen3.5 122B-A10B NVFP4, the new `gen` model (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${QWEN35_PORT:-8013}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/qwen35-vl/compose.yaml
# qwen35-vl — Qwen3.5-9B vision-language model (FP8) on ana-ml2.
#
# Co-located on GPU 1 with the granite summarizer + embed/rerank/reward trio
# (GPU 0 is deliberately kept free for hot-reloading large models). Serves on
# :8007, fronted by the LiteLLM gateway as `qwen3.5-9b-fp8`.
#
# WHY A PINNED NIGHTLY DIGEST (not :latest): vLLM :latest (v0.19.1) quantizes
# the Qwen3.5-VL *vision tower* under --quantization fp8, producing garbage
# vision output (the language model is unaffected — it answers text fine but
# "sees" noise). The nightly correctly excludes the vision tower, so vision
# works while the LM still gets the FP8 throughput/VRAM win. We pin the exact
# nightly digest for reproducibility — a moving :nightly tag would silently
# change the engine. WATCH: once the vision-FP8 exclusion lands in a stable
# release, re-pin to :latest and drop this note.
#
# WHY util 0.40 (not the trio's tiny values): the model needs ~34 GB just to
# start at 32k context (FP8 weights + BF16 vision tower + graph capture + 32k
# profiling). On shared GPU 1 (prod uses ~46 GB, ~48 GB free) this vLLM build
# requires free >= util*total, capping util at ~0.51 here; 0.40 (~38 GB) sits
# above the ~34 GB floor with ~10 GB card headroom.
#
# All tunables live in .env — edit that, not this file.
name: qwen35-vl
services:
vllm-qwen35:
image: ${QWEN_IMAGE}
container_name: ${QWEN_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${QWEN_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${QWEN_MODEL}
- --served-model-name
- ${QWEN_SERVED_NAME}
- --quantization
- fp8
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${QWEN_GPU_MEM_UTIL}
- --max-model-len
- ${QWEN_MAX_MODEL_LEN}
- --dtype
- auto
# Prefix caching pinned ON (the nightly defaults it OFF). Free win for the
# text-chat path; marginal for vision (each image is a distinct prefix).
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${QWEN_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Qwen3.5-9B VL (FP8)
- homepage.icon=mdi-image-search
- homepage.description=Qwen3.5-9B vision-language (FP8) via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/qwen36-27b-aeon/compose.yaml
# qwen36-27b-aeon — AEON-7/Qwen3.6-27B-AEON-Ultimate-Uncensored on ana-ml2 GPU 0,
# REPLACING qwopus3.5-122b as the `gen` model (operator 2026-07-05).
#
# Dense 27B, qwen3_5 GDN-hybrid arch (full-attn + Gated DeltaNet SSM) — same family
# as qwopus, VISION-INTACT (Qwen3_5ForConditionalGeneration, vision tower preserved
# at bf16), abliterated (abliterix v1.4, 0/100 refusals), native MTP head grafted,
# Apache-2.0, 131K default ctx. Served NVFP4 (ModelOpt) on Blackwell's FP4 cores.
#
# TWO CO-LOCATED INSTANCES on GPU 0 (operator wants both behaviours resident at once;
# MTP is a serve-time config, NOT per-request, so one endpoint can't do both):
# vllm-aeon-gen (:8015, served qwen3.6-27b-aeon) — MTP OFF, general/concurrent
# serve. Backs gateway gen / gen-reasoning / summarizer-large.
# vllm-aeon-rp (:8016, served qwen3.6-27b-aeon-rp) — native MTP ON (qwen3_5_mtp
# n=3), low-concurrency single-seat RP. Backs gateway char-rp.
# Why MTP off for the general serve: measured on qwopus, MTP helps single-stream
# (+12% N=1) but HURTS moderate concurrency (-15..-20% N=4) and silently drops
# min_p/logit_bias — wrong for a shared multi-consumer endpoint. Right only for a
# dedicated single-stream seat (the RP one). [[reference_gen_qwopus_122b]]
#
# VRAM budget (2 weight copies, no sharing): full NVFP4 = 27GB ea. gen util 0.45
# (~43GB) + rp util 0.40 (~38GB) = ~81GB / 96GB, ~15GB margin. depends_on:
# service_healthy sequences gen-first so the util reservation doesn't race → OOM.
# If margin bites at warmup (vision-encoder + big-vocab sampler warmup, cf. qwen36-vl),
# point AEON_RP_MODEL at the 21GB XS variant (frees ~6GB) via .env — no compose edit.
#
# NVFP4 is ModelOpt format → --quantization modelopt (vLLM also auto-detects; explicit
# is belt-and-suspenders). --mamba-cache-dtype float32 for the GDN/SSM state (AEON
# deploy guide + vLLM recipe Mamba-cache note). Tool-calling qwen3_coder + reasoning
# qwen3 (per the AEON card), same as qwopus.
#
# ⚠️ DEPLOYABILITY — load-test before trusting: multimodal + ModelOpt-NVFP4 on THIS
# brand-new arch, and native qwen3_5_mtp spec-decode, are unproven on our stock vLLM
# image. If stock can't serve it, the AEON patched image (ghcr.io/aeon-7/aeon-vllm-
# ultimate, PRs #41703/#40898) is the fallback — but that's really for DFlash; native
# MTP + base inference should ride stock >= 0.23.0. Set AEON_IMAGE in .env.
#
# REVERT: `docker compose down` here + `docker compose up -d` the qwopus3.5-122b stack
# (still staged) + revert the litellm gen/gen-reasoning/summarizer-large records.
# All tunables in .env — edit that, not this file.
name: qwen36-27b-aeon
services:
# ── General serve — MTP OFF, concurrent. gen / gen-reasoning / summarizer-large. ──
vllm-aeon-gen:
# Per-service image so gen can stay pinned to a known-good vLLM while rp tests a new one.
image: ${AEON_GEN_IMAGE:-vllm/vllm-openai:latest}
container_name: ${AEON_GEN_CONTAINER_NAME:-vllm-aeon-gen}
restart: unless-stopped
ipc: host
ports:
- "${AEON_GEN_PORT:-8015}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
- ${AEON_GEN_MODEL:-/tank/aimodels/qwen36-27b-aeon-nvfp4}:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
# Reclaims PyTorch reserved-but-unallocated fragmentation so the co-located
# util split doesn't strand VRAM (same knob qwopus needed for the MoE workspace).
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
command:
- /model
# TWO served-names: base + a `-thinking` alias. LiteLLM keys deployments by
# (model, api_base), so gen and gen-reasoning MUST use distinct model names or a
# thinking-off request mutates the shared litellm_params and clobbers the other's
# enable_thinking (the shared-config-mutation footgun). gen-reasoning routes to the
# `-thinking` name; gen/summarizer-large route to the base name.
- --served-model-name
- ${AEON_GEN_SERVED_NAME:-qwen3.6-27b-aeon}
- ${AEON_GEN_SERVED_NAME_THINK:-qwen3.6-27b-aeon-thinking}
- --host
- 0.0.0.0
- --port
- "8000"
- --quantization
- ${AEON_GEN_QUANT:-modelopt}
- --gpu-memory-utilization
- ${AEON_GEN_GPU_MEM_UTIL:-0.45}
- --max-model-len
- ${AEON_GEN_MAX_MODEL_LEN:-131072}
# Keep concurrency modest: big-vocab sampler warmup allocates a large tensor
# (qwen36-vl OOM'd at the default 1024 on a shared GPU). 16 is ample here.
- --max-num-seqs
- ${AEON_GEN_MAX_NUM_SEQS:-16}
- --max-num-batched-tokens
- "16384"
- --trust-remote-code
- --dtype
- auto
# GDN/SSM (Gated DeltaNet) state cache — float32 per the AEON deploy guide.
- --mamba-cache-dtype
- float32
- --kv-cache-dtype
- ${AEON_GEN_KV_CACHE_DTYPE:-fp8}
- --enable-prefix-caching
- --enable-chunked-prefill
- --limit-mm-per-prompt
- '{"image": 4}'
- --reasoning-parser
- ${AEON_GEN_REASONING_PARSER:-qwen3}
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${AEON_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=Qwen3.6-27B AEON (NVFP4, vision) — gen
- homepage.icon=mdi-creation
- homepage.description=Uncensored Qwen3.6-27B multimodal NVFP4, the `gen` model (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${AEON_GEN_PORT:-8015}/docs
# ── RP seat — native MTP ON, low concurrency, single-seat. char-rp. ──
vllm-aeon-rp:
image: ${AEON_RP_IMAGE:-vllm/vllm-openai:latest}
container_name: ${AEON_RP_CONTAINER_NAME:-vllm-aeon-rp}
restart: unless-stopped
ipc: host
# Sequence AFTER the general serve is healthy so the two util reservations on the
# shared GPU don't race into an OOM (gen reserves its 0.45 first, then rp its 0.40).
depends_on:
vllm-aeon-gen:
condition: service_healthy
ports:
- "${AEON_RP_PORT:-8016}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
# Point at the XS (21GB) variant via .env to buy ~6GB co-location margin.
- ${AEON_RP_MODEL:-/tank/aimodels/qwen36-27b-aeon-nvfp4}:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
command:
- /model
# base + `-thinking` alias (see the gen note): char-rp -> base, char-rp-reasoning -> -thinking.
- --served-model-name
- ${AEON_RP_SERVED_NAME:-qwen3.6-27b-aeon-rp}
- ${AEON_RP_SERVED_NAME_THINK:-qwen3.6-27b-aeon-rp-thinking}
- --host
- 0.0.0.0
- --port
- "8000"
- --quantization
- ${AEON_RP_QUANT:-modelopt}
- --gpu-memory-utilization
- ${AEON_RP_GPU_MEM_UTIL:-0.40}
- --max-model-len
- ${AEON_RP_MAX_MODEL_LEN:-65536}
# Single-seat: low concurrency keeps warmup + KV small so it fits alongside gen.
- --max-num-seqs
- ${AEON_RP_MAX_NUM_SEQS:-2}
- --trust-remote-code
- --dtype
- auto
- --mamba-cache-dtype
- float32
- --kv-cache-dtype
- ${AEON_RP_KV_CACHE_DTYPE:-fp8}
- --enable-prefix-caching
- --limit-mm-per-prompt
- '{"image": 4}'
- --reasoning-parser
- ${AEON_RP_REASONING_PARSER:-qwen3}
# Native MTP speculative decode (the grafted head). n=3 per the AEON card's
# measured accept length (~3.3/3). MTP is why this seat exists separately.
- --speculative-config
- '{"method": "${AEON_RP_SPEC_METHOD:-qwen3_5_mtp}", "num_speculative_tokens": ${AEON_RP_SPEC_TOKENS:-3}}'
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${AEON_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI - Dormant
- homepage.name=Qwen3.6-27B AEON RP (NVFP4 + MTP) — char-rp
- homepage.icon=mdi-drama-masks
- homepage.description=Uncensored Qwen3.6-27B, native MTP single-seat RP (char-rp), ana-ml2 GPU 0
- homepage.href=http://10.250.50.54:${AEON_RP_PORT:-8016}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/qwen36-vl/compose.yaml
# qwen36-vl — Qwen3.6-35B-A3B vision-language MoE (official NVFP4) on ana-ml2.
#
# Replaces the qwen35-vl stack (Qwen3.5-9B) 2026-06-14. Co-located on GPU 1 with
# the granite summarizer + embed/rerank/reward trio. Serves on :8007.
#
# WHY NVFP4 now (swapped FROM FP8 2026-06-15): the nvidia ModelOpt NVFP4 MoE that
# was BROKEN on vLLM 0.19.1/0.22.0 (#44081, lm_head.input_scale) loads clean on
# 0.23.0 — the ModelOpt lm_head fix landed. So we cut FP8→NVFP4: ~20.4 GiB weights
# vs FP8's ~34 GiB (~40% lighter, ~13 GB reclaimed on GPU 1), faster single-stream
# on Blackwell's FP4 tensor cores, and the freed room funds fp16 KV + a granite
# context restore (see the rebalance note below). The NVFP4 checkpoint preserves
# the vision tower (ModelOpt leaves it high-precision) — VALIDATED by comfy-dev's
# real anatomy-judge A/B on 16 prod images: PASS, holds the load-bearing
# discrimination (gross-deformity reject + clean-pass), only shuffles already-
# unreliable sub-ceiling borderline-hand calls. brokkr's text/speed arm: parity
# except a minor multi-step chained-numeric-reasoning slip (W4A4 tell) — doesn't
# bite the vision-judge role; flag for any gateway consumer doing chained math.
# Requires vLLM >= 0.23.0 (pinned by digest in .env). NO --quantization flag
# (vLLM auto-detects the checkpoint's NVFP4).
#
# NAMING: served ONLY as its TRUE name `qwen3.6-35b-a3b`. A model is never aliased
# under a prior model's name — a caller asking for `qwen3.5-9b-fp8` (a 9B dense)
# must NOT be silently handed this 35B-A3B MoE; that's a downstream-confusion
# footgun. The legacy `qwen3.5-9b-fp8` name is RETIRED. Consumers (Arbo's vision
# hero-judge, stacks/arbo v0.11.3+) migrate to `qwen3.6-35b-a3b` — they 404 on the
# old name until they repoint, which is the correct loud signal (notified 2026-06-14).
#
# GPU-1 REBALANCE (2026-06-15, pinned): the NVFP4 swap freed ~13 GB, redistributed —
# qwen36 NVFP4: util 0.46→0.32 (~31 GB: 20.4 GB weights + fp16 KV + graph).
# fp16 KV (we DROPPED --kv-cache-dtype fp8) — the freed room buys back full-
# precision KV; hybrid attn (10/40 full-attn) keeps even fp16 KV affordable.
# granite: RESTORED 0.24→0.34, max-len 65536→131072 (gives back the context
# sacrificed for FP8 qwen — the FP8-vs-maxed-granite tradeoff is now undone).
# trio (embed/rerank/reward) unchanged at floor.
# Total GPU-1 util ~0.82 → ~17 GB headroom (was a tight ~5 GB).
#
# THINKING TOGGLE: this is ONE hybrid checkpoint (not separate Instruct/Thinking
# downloads) with a Qwen3-style per-request `enable_thinking` switch. The chat
# template defaults thinking ON (`<think>\n`); passing
# `chat_template_kwargs={"enable_thinking":false}` emits the empty
# `<think>\n\n</think>\n\n` block (no reasoning). We run --reasoning-parser qwen3
# (model-matched — its vLLM docstring describes THIS checkpoint) so ONE endpoint
# serves BOTH modes cleanly: thinking-ON splits <think>…</think> into
# reasoning_content; thinking-OFF routes everything to content. The gateway
# selects the mode per model_name (stacks/litellm/conf/config.yaml):
# qwen3.6-35b-a3b → enable_thinking:false (non-thinking DEFAULT)
# qwen3.6-35b-a3b-thinking → enable_thinking:true (opt-in reasoning)
#
# All tunables live in .env — edit that, not this file.
name: qwen36-vl
services:
vllm-qwen36:
image: ${QWEN_IMAGE}
container_name: ${QWEN_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${QWEN_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${QWEN_MODEL}
# Pre-quantized NVFP4 (ModelOpt) checkpoint → NO --quantization (vLLM auto-
# detects; the vision tower is left high-precision by the producer).
- --served-model-name
- qwen3.6-35b-a3b
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${QWEN_GPU_MEM_UTIL}
- --max-model-len
- ${QWEN_MAX_MODEL_LEN}
# Cap concurrency: vLLM warms the sampler with max_num_seqs dummy requests,
# and this model's 248K vocab makes that warmup tensor huge — the default
# 1024 OOMs on a shared GPU even though weights+KV fit. 32 is ample for a
# vision endpoint (the summarizer carries the concurrency, not this).
- --max-num-seqs
- ${QWEN_MAX_NUM_SEQS}
# fp16 KV (no --kv-cache-dtype): the NVFP4 swap freed enough room to run
# full-precision KV — better than the fp8 KV the FP8 build needed to fit.
- --trust-remote-code
- --dtype
- auto
- --enable-prefix-caching
# Model-matched reasoning parser for the hybrid thinking toggle (see header).
# Splits <think>…</think> into reasoning_content when thinking is ON; routes
# all output to content when the empty think-block signals thinking OFF — so
# this single :8007 endpoint serves both the non-thinking default and the
# qwen3.6-35b-a3b-thinking gateway variant.
- --reasoning-parser
- qwen3
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${QWEN_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=Qwen3.6-35B-A3B VL (NVFP4)
- homepage.icon=mdi-image-search
- homepage.description=Qwen3.6-35B-A3B vision-language MoE (NVFP4) via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${QWEN_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/qwen-image-bench/compose.yaml
# qwen-image-bench — flukethoughts/Qwen-Image-Bench-NVFP4 on ana-ml2 GPU 1,
# REPLACING qwen3.6-35b-a3b (operator 2026-06-19). Qwen's text-to-image quality
# JUDGE model (vision-language, NVFP4 weights / vision tower bf16). NOT generative —
# it scores T2I outputs on 5 dims (overall quality, prompt match, aesthetic, LoRA
# activation, confidence).
#
# Arch: Qwen3_5ForConditionalGeneration (dense Qwen3.5 hybrid SSM+attn + vision),
# ~17B / ~20GB NVFP4. VISION-INTACT → served as multimodal; NO text-only patch
# (unlike the qwen3.5-122b gen model, which had text-only weights). vLLM
# production-validated per the model card.
#
# ⚠️ qwen3.6-35b-a3b was arbo's hero-judge (comfy-dev consumer). Downing it breaks
# arbo's judging until comfy-dev repoints to qwen-image-bench (different I/O — a
# 5-dim verdict vs a general VL judge). comfy-dev notified.
#
# DISPLACEMENT: GPU 1 is shared (granite/selene/embed/rerank/reward). qwen3.6 used
# util 0.34 (~33GB); down it first, then this fits at util ~0.22 (~21GB). REVERT =
# down this, `up -d` the qwen36 stack.
#
# Tunables in .env.
name: qwen-image-bench
services:
vllm-qwen-image-bench:
image: ${QIB_IMAGE:-vllm/vllm-openai:latest}
container_name: ${QIB_CONTAINER_NAME:-vllm-qwen-image-bench}
restart: unless-stopped
ipc: host
ports:
- "${QIB_PORT:-8014}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
- /tank/aimodels/qwen-image-bench-nvfp4:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
command:
- /model
- --served-model-name
- qwen-image-bench
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${QIB_GPU_MEM_UTIL:-0.32}
- --max-model-len
- ${QIB_MAX_MODEL_LEN:-32768}
- --max-num-seqs
- ${QIB_MAX_NUM_SEQS:-8}
- --trust-remote-code
- --dtype
- auto
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${QIB_GPU_ID:-1}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 600s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=Qwen-Image-Bench (T2I judge, NVFP4)
- homepage.icon=mdi-image-check
- homepage.description=Qwen text-to-image quality judge (NVFP4, vision-intact) on ana-ml2 GPU 1
- homepage.href=http://10.250.50.54:${QIB_PORT:-8014}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/qwopus3.5-122b/compose.yaml
# qwopus3.5-122b — OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4
# on ana-ml2 GPU 0, REPLACING the bjk110 text-only qwen3.5-122b as the `gen` model
# (operator 2026-06-19: "already ablated, already quanted, vision tower intact").
#
# Qwen3.5-122B-A10B MoE, Kimi-K2.6-distilled + abliterated, NVFP4 — and crucially
# VISION-INTACT (Qwen3_5MoeForConditionalGeneration + vision_config). So it serves as
# plain MULTIMODAL (no text-only patch, unlike the bjk110 checkpoint which had its
# vision weights stripped). vLLM carries the arch natively.
#
# Served under --served-model-name qwen3.5-122-a10b so the existing litellm records
# (gen / gen-reasoning / qwen3.5-122-a10b[-reasoning] / qwen-large[-reasoning]) route
# here UNCHANGED — the operator's "replace those records with this model". The thinking
# split (chat_template_kwargs.enable_thinking) + tool-calling (qwen3_coder — the
# OpenYourMind card's specified parser for this checkpoint's XML tool calls).
#
# REVERT: down this; the bjk110 qwen3.5-122b stack is still staged.
# Tunables in .env.
name: qwopus3.5-122b
services:
vllm-qwopus35-122b:
image: ${QWOPUS_IMAGE:-vllm/vllm-openai:latest}
container_name: ${QWOPUS_CONTAINER_NAME:-vllm-qwopus35-122b}
restart: unless-stopped
ipc: host
ports:
- "${QWOPUS_PORT:-8013}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
- /tank/aimodels/qwopus3.5-122b-nvfp4:/model:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
# Reclaims PyTorch's reserved-but-unallocated fragmentation (4.2GB was stranded at
# util 0.96, starving the FusedMoE workspace → OOM by 0.1GB). Lets the 3.09GB MoE
# workspace allocate cleanly. Same knob the bjk110 qwen3.5-122b stack ran.
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
command:
- /model
- --served-model-name
- ${QWOPUS_SERVED_NAME:-qwen3.5-122-a10b}
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${QWOPUS_GPU_MEM_UTIL:-0.92}
- --max-model-len
- ${QWOPUS_MAX_MODEL_LEN:-131072}
- --max-num-seqs
- ${QWOPUS_MAX_NUM_SEQS:-8}
- --max-num-batched-tokens
- "32768"
- --trust-remote-code
- --dtype
- auto
- --enable-prefix-caching
- --enable-chunked-prefill
# FULL 256K context on the STABLE image. fp8 KV (near-lossless) measured an 11.8GB
# pool = 934,600 tokens = 3.5x concurrency at the full 262144 window. CUDA graphs ON
# (no --enforce-eager) for decode tok/s. BINDING LIMIT = the FusedMoE transient
# workspace (3.09GB, allocated OUTSIDE vLLM's budget into free VRAM): at util 0.96
# only 2.99GB was free → OOM by 0.1GB, worsened by 4.2GB PyTorch fragmentation.
# FIX = expandable_segments (env above, reclaims the fragmentation) + util 0.95 for
# margin. The card can't go to 0 free — this workspace is the floor. video kept
# ENABLED (operator wants it; banked at util 0.95 with headroom) — the video encoder
# profiling eats into the budget so KV concurrency drops some, but stays well above 2x.
- --kv-cache-dtype
- ${QWOPUS_KV_CACHE_DTYPE:-fp8}
- --limit-mm-per-prompt
- '{"image": 2, "video": 1}'
# reasoning split + tool-calling. The OpenYourMind card specifies qwen3_coder
# as the tool-call parser for this checkpoint (Qwen3.5 XML tool-call format).
- --reasoning-parser
- qwen3
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${QWOPUS_GPU_ID:-0}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 900s
networks:
- tnet
labels:
- homepage.group=AI - Dormant
- homepage.name=Qwopus3.5-122B-A10B (abliterated NVFP4, vision)
- homepage.icon=mdi-creation
- homepage.description=Kimi-distilled abliterated Qwen3.5-122B-A10B NVFP4, vision-intact, the `gen` model (ana-ml2 GPU 0)
- homepage.href=http://10.250.50.54:${QWOPUS_PORT:-8013}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/selene/compose.yaml
# selene — AtlaAI Selene 1 Mini (Llama 3.1 8B) judge/eval model on ana-ml2 GPU 1.
#
# Restores the judge that went offline when llama-swap was downed (it was the
# Q6_K GGUF `selene-1-mini-8b` in the llama-swap zoo). Re-served on vLLM at the
# operator's request, FP8 (NVFP4 had no pre-made checkpoint and W4A4 is too
# aggressive for a precision judge validated at Q6_K — FP8 ≥ Q6_K fidelity).
#
# FP8 = vLLM DYNAMIC --quantization fp8 (W8A8) of the bf16 AtlaAI checkpoint —
# no offline quant needed, near-lossless, and Selene is text-only Llama 3.1 so
# there's NO vision tower for dynamic fp8 to noise-quantize (the qwen35-VL
# footgun doesn't apply here). ~8 GiB weights on GPU 1's headroom.
#
# Co-tenant on GPU 1 with qwen36 (NVFP4) + granite + embed/rerank/reward. Sized
# to fit the ~24 GB headroom while leaving GPU 1 a safe buffer (see .env).
# Served ONLY as `selene-1-mini-8b` (the name its consumers know). All tunables
# live in .env.
name: selene
services:
vllm-selene:
image: ${SELENE_IMAGE}
container_name: ${SELENE_CONTAINER_NAME}
restart: unless-stopped
ipc: host
ports:
- "${SELENE_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${SELENE_MODEL}
# Dynamic FP8 (W8A8) from the bf16 checkpoint — no pre-quant needed.
- --quantization
- fp8
- --served-model-name
- selene-1-mini-8b
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${SELENE_GPU_MEM_UTIL}
- --max-model-len
- ${SELENE_MAX_MODEL_LEN}
- --max-num-seqs
- ${SELENE_MAX_NUM_SEQS}
# fp8 KV — matches the judge's old q8 KV posture + keeps the pool compact
# on the shared card.
- --kv-cache-dtype
- fp8
- --dtype
- auto
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${SELENE_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=Selene 1 Mini 8B (judge, FP8)
- homepage.icon=mdi-gavel
- homepage.description=AtlaAI Selene 1 Mini Llama-3.1-8B judge (FP8) via vLLM (ana-ml2 GPU1)
- homepage.href=http://10.250.50.54:${SELENE_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/vibevoice/compose.yaml
services:
vibevoice:
container_name: vibevoice
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities:
- gpu
ports:
- 8745:8745
volumes:
- /tank/vibevoice/hf:/root/.cache/huggingface
- /tank/vibevoice/voices:/app/voices
- /tank/vibevoice/state:/var/lib/eworker
environment:
- ENABLE_1_5B=true
- ENABLE_LARGE=true
- AUTH_REQUIRED=true
- CORS_ENABLED=true
- ALLOWED_ORIGINS=*
image: eworkerinc/vibevoice:latest
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=VibeVoice
- homepage.icon=mdi-chat
- homepage.description=EWorkerStudio VibeVoice
- homepage.href=http://10.250.50.54:8745
restart: unless-stopped
networks:
tnet:
name: traefik-net
external: true
>>> /opt/docker/compose/vllm/compose.yaml
# vLLM — Qwen3 Embedding + Reranker + Skywork Reward-V2 classifier.
#
# Originally created to replace the unmaintained Infinity stack (embed +
# rerank); generalized 2026-05-13 to host any vLLM-served model on ana-ml2,
# starting with the Skywork-Reward-V2-Llama-3.1-8B reward classifier
# (AWQ-quantized locally, served from /tank/aimodels/llm/).
#
# vLLM runs one model per process, so this stack brings up three containers
# sharing a single GPU:
#
# vllm-embed — Qwen3-Embedding served as an OpenAI /v1/embeddings server
# vllm-rerank — Qwen3-Reranker served as a /rerank + /score server
# vllm-reward — Skywork-Reward-V2-Llama-3.1-8B-AWQ served as a /classify scorer
#
# The reranker is a causal-LM checkpoint; --hf-overrides re-maps it to
# Qwen3ForSequenceClassification so vLLM's reranking endpoints work and the
# model only emits two class logits (no/yes) instead of the full 151k vocab.
#
# All tunables live in .env — edit that, not this file.
#
# Pre-download models to avoid first-run delay:
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Embedding-0.6B
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Reranker-0.6B
#
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on ana-ml2 and is
# bind-mounted into the reward service at /local-models. Not from HF Hub.
services:
vllm-embed:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-embed
restart: unless-stopped
ipc: host
ports:
- "${EMBED_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${EMBED_MODEL}
- --served-model-name
- ${EMBED_MODEL}
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${EMBED_GPU_MEM_UTIL}
- --max-model-len
- ${EMBED_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Embed (Qwen3)
- homepage.icon=mdi-vector-arrange-below
- homepage.description=Qwen3 Embedding via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${EMBED_PORT}/docs
vllm-rerank:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-rerank
restart: unless-stopped
ipc: host
ports:
- "${RERANK_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
- ${RERANK_MODEL}
- --served-model-name
- ${RERANK_MODEL}
- --runner
- pooling
- --hf-overrides
- '{"architectures":["Qwen3ForSequenceClassification"],"classifier_from_token":["no","yes"],"is_original_qwen3_reranker":true}'
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${RERANK_GPU_MEM_UTIL}
- --max-model-len
- ${RERANK_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Rerank (Qwen3)
- homepage.icon=mdi-sort-variant
- homepage.description=Qwen3 Reranker via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${RERANK_PORT}/docs
vllm-reward:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-reward
restart: unless-stopped
ipc: host
ports:
- "${REWARD_PORT}:8000"
volumes:
# AWQ output lives in the legacy llama-swap models tree, not the HF cache
# — bind-mount the LLM models dir read-only so the reward service can
# load it as a local-path HF-format model.
- /tank/aimodels/llm:/local-models:ro
environment:
- VLLM_API_KEY=${API_KEY:-}
command:
- /local-models/Skywork-Reward-V2-Llama-3.1-8B-AWQ
- --served-model-name
- Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
# vLLM 0.19.1 deprecated --task in favor of --runner. The model's
# config.json declares `LlamaForSequenceClassification` so the
# pooling runner uses it as a classifier (single-label reward score)
# without needing an explicit task flag.
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${REWARD_GPU_MEM_UTIL}
- --max-model-len
- ${REWARD_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 240s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Reward (Skywork)
- homepage.icon=mdi-scale-balance
- homepage.description=Skywork-Reward-V2 8B classifier via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${REWARD_PORT}/docs
# Phi-4-mini (FP8) — summarizer + "dreaming" agent. Supersedes the
# llama-swap granite-4-small pin. Generative chat model (OpenAI
# /v1/chat/completions), so NO --runner pooling. FP8 on RTX 6000 Ada
# (cc 8.9): near-lossless, ~1.2x, ~6 GB.
vllm-granite:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-granite
restart: unless-stopped
ipc: host
ports:
- "${GRANITE_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- HUGGING_FACE_HUB_TOKEN=${HF_TOKEN:-}
- VLLM_API_KEY=${API_KEY:-}
command:
# Production summarizer (replaced phi4-mini 2026-06-05). Default = official
# IBM pre-quantized FP8 (compressed-tensors), loaded directly; FP8 is native
# on the RTX 6000 Ada (cc 8.9). Fallback to vLLM-native dynamic FP8 from
# BF16: GRANITE_MODEL=ibm-granite/granite-4.1-8b + GRANITE_QUANT=fp8.
- ${GRANITE_MODEL}
- --served-model-name
- ${GRANITE_SERVED_NAME}
- --quantization
- ${GRANITE_QUANT}
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${GRANITE_GPU_MEM_UTIL}
- --max-model-len
- ${GRANITE_MAX_MODEL_LEN}
- --max-num-seqs
- ${GRANITE_MAX_NUM_SEQS}
- --dtype
- auto
# CUDA graphs ENABLED (no --enforce-eager) for decode throughput. Made
# room 2026-06-05 by right-sizing the embed/rerank/reward trio's KV pools
# (they were over-provisioned at 5.9x/2.0x/3.9x concurrency); GPU 1 now has
# ~17 GB free after granite, so graph-capture buffers fit. If the trio
# ever grows back, granite may need --enforce-eager again on this card.
# FP8 KV cache — halves KV memory; near-lossless on Ada (cc 8.9).
- --kv-cache-dtype
- ${GRANITE_KV_CACHE_DTYPE}
# Prefix caching pinned EXPLICIT (vLLM v1 defaults it on, but pin so a
# version flip can't silently disable it). Benched 2026-06-13: ~6.5x faster
# TTFT (45ms vs 292ms) on a shared ~4.5k-token summarizer template; soft/
# evictable KV, neutral when prefixes don't repeat — pure win for granite.
- --enable-prefix-caching
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GRANITE_GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 180s
networks:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=vLLM Granite 4.1 8B (summarizer)
- homepage.icon=mdi-text-box-outline
- homepage.description=Granite 4.1 8B FP8 via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${GRANITE_PORT}/docs
networks:
tnet:
name: traefik-net
external: true
===== CONFIG LAYOUT (/opt/docker/conf/ — top 200 entries) =====
/opt/docker/conf
/opt/docker/conf/llama-swap
/opt/docker/conf/llama-swap/config.yaml
/opt/docker/conf/vllm
===== LISTENING PORTS =====
0.0.0.0:111
0.0.0.0:22
0.0.0.0:5001
0.0.0.0:7007
0.0.0.0:8001
0.0.0.0:8002
0.0.0.0:8003
0.0.0.0:8004
0.0.0.0:8011
0.0.0.0:8015
0.0.0.0:8016
0.0.0.0:8018
[::]:111
[::]:22
*:2375
[::]:5001
[::]:8001
[::]:8002
[::]:8003
[::]:8004
[::]:8011
[::]:8015
[::]:8016
[::]:8018
===== MODEL / HUGGINGFACE CACHES =====
/tank/aimodels/huggingface (604G)
hub entries:
CACHEDIR.TAG
datasets--HuggingFaceH4--ultrachat_200k
datasets--mlabonne--harmful_behaviors
datasets--mlabonne--harmless_alpaca
datasets--neuralmagic--calibration
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
datasets--wikitext
models--AtlaAI--Selene-1-Mini-Llama-3.1-8B
models--AxionML--Qwen3.5-9B-NVFP4
models--bartowski--Meta-Llama-3.1-8B-Instruct-GGUF
models--bartowski--NousResearch_Hermes-4-14B-GGUF
models--bartowski--TheDrummer_GLM-Steam-106B-A12B-v1-GGUF
models--bartowski--TheDrummer_Skyfall-31B-v4-GGUF
models--BeaverAI--Artemis-31B-v1i-GGUF
models--BeaverAI--Skyfall-R1-31B-v4a-GGUF
models--DavidAU--Qwen3.6-27B-Heretic2-Uncensored-Finetune-Thinking
models--drawais--Granite-4.1-30B-NVFP4
models--ibm-granite--granite-4.0-h-small-GGUF
models--ibm-granite--granite-4.0-h-tiny-GGUF
models--ibm-granite--granite-4.0-micro-GGUF
models--ibm-granite--granite-4.1-8b
models--ibm-granite--granite-4.1-8b-fp8
models--llmfan46--Qwen3.6-35B-A3B-uncensored-heretic-GGUF
models--microsoft--Phi-4-mini-instruct
models--mistralai--Mistral-Small-4-119B-2603-NVFP4
models--mradermacher--Daredevil-8B-abliterated-dpomix-GGUF
models--mradermacher--Qwen3-30B-A3B-abliterated-erotic-i1-GGUF
models--mradermacher--Qwen3.6-35B-A3B-abliterated-i1-GGUF
models--mradermacher--Selene-1-Mini-Llama-3.1-8B-i1-GGUF
models--murilonwt--granite-4.1-8b-NVFP4
/tank/aimodels/llm (1.1T)
/home/lkraven/.cache/huggingface (18G)
hub entries:
CACHEDIR.TAG
datasets--HuggingFaceH4--ultrachat_200k
datasets--Salesforce--wikitext
datasets--Skywork--Skywork-Reward-Preference-80K-v0.2
datasets--wikitext
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP
models--AEON-7--Qwen3.6-27B-AEON-Ultimate-Uncensored-Multimodal-NVFP4-MTP-XS
models--Astralyra--bge-reranker-large-Q8_0-GGUF
models--bicro--qwen3.5-abliterated-vision-merged
models--bjk110--Qwen3.5-122B-A10B-abliterated-NVFP4
models--darkc0de--Mistral-Small-4-119B-2603-heretic
models--DavidAU--Qwen3.6-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-Thinking
models--flukethoughts--Qwen-Image-Bench-NVFP4
models--ggml-org--embeddinggemma-300M-GGUF
models--ggml-org--Qwen3-Reranker-0.6B-Q8_0-GGUF
models--huihui-ai--Huihui-Qwen3.5-122B-A10B-abliterated
models--jinaai--jina-reranker-v3-GGUF
models--klnstpr--bge-reranker-v2-m3-Q8_0-GGUF
models--minhtd14--jina-reranker-v2-base-multilingual-Q8_0-GGUF
models--mistralai--Mistral-Medium-3.5-128B-EAGLE
models--mistralai--Mistral-Small-4-119B-2603-NVFP4
models--Mungert--Qwen3-Reranker-0.6B-GGUF
models--OpenYourMind--Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4
models--OpenYourMind--Qwopus3.5-122B-A10B-Kimi-K2.6-destill-healed-abliterated
models--Qwen--Qwen3-Embedding-0.6B-GGUF
models--Qwen--Qwen3-Reranker-0.6B
models--RecViking--Mistral-Medium-3.5-128B-NVFP4
models--robbatt--Qwen3.6-40B-Deckard-NVFP4
models--Skywork--Skywork-Reward-V2-Llama-3.1-8B
===== DOCKER-ADJACENT SYSTEMD SERVICES =====
containerd.service running
docker.service running
nvidia-persistenced.service running
===== DONE =====
Paste the above back into the chat, or pass a path as argv[1] to save.