20e796cf6b
Replaces the bjk110 text-only qwen3.5-122b as the `gen` model on ana-ml2 GPU 0. OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 — Kimi- distilled, abliterated, NVFP4, and crucially VISION-INTACT (serves as plain multimodal, no text-only patch). Served as qwen3.5-122-a10b so the litellm gen / gen-reasoning / qwen-large records route here unchanged. Tuned for full native context on the 96GB Blackwell: - stable vLLM image + fp8 KV → 11GB pool = 870,014 tokens = 3.32x concurrency at the full 262144 (256K) window. Nightly+turboquant-4bit was unnecessary. - CUDA graphs ON (no --enforce-eager) → 92.7 tok/s warm single-stream. - util 0.95 + PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True — 0.96 OOM'd by 0.1GB on the 3.09GB FusedMoE transient workspace (the hard floor; defrag reclaims the 4.2GB fragmentation, 0.95 adds margin). - max-num-seqs 16 (short reqs fan out ~16x32k; 256K reqs pool-limit to 3.32x). - text + image + video all enabled; tool-calling via qwen3_coder (XML), verified.