20e796cf6b
Replaces the bjk110 text-only qwen3.5-122b as the `gen` model on ana-ml2 GPU 0. OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 — Kimi- distilled, abliterated, NVFP4, and crucially VISION-INTACT (serves as plain multimodal, no text-only patch). Served as qwen3.5-122-a10b so the litellm gen / gen-reasoning / qwen-large records route here unchanged. Tuned for full native context on the 96GB Blackwell: - stable vLLM image + fp8 KV → 11GB pool = 870,014 tokens = 3.32x concurrency at the full 262144 (256K) window. Nightly+turboquant-4bit was unnecessary. - CUDA graphs ON (no --enforce-eager) → 92.7 tok/s warm single-stream. - util 0.95 + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True — 0.96 OOM'd by 0.1GB on the 3.09GB FusedMoE transient workspace (the hard floor; defrag reclaims the 4.2GB fragmentation, 0.95 adds margin). - max-num-seqs 16 (short reqs fan out ~16x32k; 256K reqs pool-limit to 3.32x). - text + image + video all enabled; tool-calling via qwen3_coder (XML), verified.
29 lines
1.4 KiB
Bash
29 lines
1.4 KiB
Bash
# qwopus3.5-122b (OpenYourMind Qwopus3.5-122B-A10B Kimi-distilled abliterated NVFP4,
|
|
# vision-intact) — ana-ml2 GPU 0 tunables. Real .env at /opt/docker/compose/qwopus3.5-122b/.env
|
|
|
|
# STABLE image + fp8 KV reaches full 256K: the KV pool already held ~222k tokens, so fp8
|
|
# (near-lossless, half the bytes/token) clears 262144 with ~2x concurrency. Nightly +
|
|
# turboquant_4bit_nc would buy ~5x concurrency at 256K but adds a 4-bit recall risk + FA2
|
|
# fallback + nightly instability — not needed for 256K itself.
|
|
QWOPUS_IMAGE=vllm/vllm-openai:latest
|
|
QWOPUS_CONTAINER_NAME=vllm-qwopus35-122b
|
|
QWOPUS_KV_CACHE_DTYPE=fp8
|
|
|
|
# Reuse :8013 (the bjk110 qwen3.5-122b port, now retired) so the litellm records route
|
|
# here unchanged. Served under qwen3.5-122-a10b (the operator's gen records).
|
|
QWOPUS_PORT=8013
|
|
QWOPUS_SERVED_NAME=qwen3.5-122-a10b
|
|
QWOPUS_GPU_ID=0
|
|
|
|
# Vision-intact NVFP4 (≈82GB incl. bf16 vision tower) on the 96GB Blackwell. CUDA graphs
|
|
# ON (no --enforce-eager) for decode throughput. util 0.95 — 0.96 OOM'd by 0.1GB on the
|
|
# 3.09GB FusedMoE transient workspace (the hard floor; the card can't reach 0 free), so
|
|
# expandable_segments (compose env) reclaims PyTorch fragmentation + 0.95 adds margin.
|
|
# max-num-seqs 16 lets short requests fan out (~16x32k); 256K requests pool-limit to ~3.5x.
|
|
QWOPUS_GPU_MEM_UTIL=0.95
|
|
QWOPUS_MAX_MODEL_LEN=262144
|
|
QWOPUS_MAX_NUM_SEQS=16
|
|
|
|
# Optional upstream vLLM API key (empty = no auth; internal net only).
|
|
API_KEY=
|