# mog-sec — pen-test seat (ana-ml2 GPU1, :8019). Copy to .env on the host. # # Values below are the configuration VALIDATED 2026-08-22: DFlash2 speculative # decoding on a newer vLLM, 420K context, 2048x2048 vision. Promoted from a # standalone experimental container after real-use testing, then re-tuned down # from 0.55/480K after that combination crashed the engine in production (see # the memory notes below — it was a headroom failure, not a context failure). # ── Image ─────────────────────────────────────────────────────────────────── # Contains DFlash2 (#52816) AND GDN spec-decode fix #53077. Verified by # ancestry, not by version string: this reports 0.26.1rc1.dev1102, which LOOKS # older than the 0.27.2rc1.dev150 it replaced — that is a setuptools_scm # tag-reachability artifact, not a downgrade. It is +259 commits, behind_by=0. # ⚠ Bumping this is a CONFIG-COMPATIBILITY event, not just a version change. # Newer transformers enforces model-config constraints older ones ignored. MOG_IMAGE=vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013 API_KEY= MOG_GPU_ID=1 MOG_CONTAINER_NAME=vllm-mog-sec MOG_PORT=8019 MOG_SERVED_NAME=mog-sec-27b MOG_SERVED_NAME_THINK=mog-sec-27b-thinking MOG_MODEL=/tank/aimodels/mog-sec-27b-nvfp4-mixed MOG_QUANT=compressed-tensors # ── Speculative decoding — DFlash2 ────────────────────────────────────────── # 2B block-diffusion drafter, drafts a whole 8-token block in one pass. Unlike # the MTP head it reads the TARGET's live hidden states (layers 5/19/33/47/61), # so it is model-agnostic — the same weights measured identically against a # different finetune (3.254 vs 3.252 accepted tok/forward, a 0.06% delta). # Measured on this seat: 2.676 -> 3.252 tok/forward, 110.5 -> 130.0 tok/s. # ⚠ The drafter is COUPLED to its target and lives in that engine's process: # the weights file is shareable across seats, the 3.85 GB of VRAM is NOT. MOG_DRAFT_MODEL=/tank/aimodels/qwen38-27b-dflash2-drafter MOG_SPEC_CONFIG={"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7} # ROLLBACK to the previous behaviour — one line, plus MOG_IMAGE above: # MOG_SPEC_CONFIG={"method": "qwen3_5_mtp", "num_speculative_tokens": 3} # ── Context and memory ────────────────────────────────────────────────────── # The model carries a complete YaRN config (rope_type yarn, factor 4.0, # original_max_position_embeddings 262144, max_position_embeddings 1000000), # so context is a KV-MEMORY choice, not a model limit. # ⚠ max-model-len must stay UNDER the KV pool or vLLM refuses to start — and # ⭐ THE POOL VARIES BOOT TO BOOT. Measured at 0.52 on two consecutive boots: # 425,663 then 453,600 tokens, a 6.6% swing. The pool is sized from FREE # memory at startup, which depends on what GPU1's other five vLLM tenants # happen to hold at that instant. 420,000 sits under the LOWER observation, # so it starts reliably in both conditions. Sizing to the higher one would # boot today and refuse tomorrow. # Concurrency ~1.0-1.08x: one request at FULL length. Shorter requests still # batch normally — vLLM schedules by tokens, not by max-model-len. MOG_MAX_MODEL_LEN=420000 # ⚠ 0.55 CRASHED IN PRODUCTION. It booted fine and then died at RUNTIME on an # ordinary ~20k-token request: torch.OutOfMemoryError allocating 614 MiB in # the GDN causal_conv1d path, with 443 MiB free. gpu-memory-utilization is # NOT a hard reservation — weights and KV are reserved, but ACTIVATION grows # on demand past the startup profile (profiled 3.62 GiB on dummy data; the # process sat ~1 GiB over its whole budget even at idle). With six vLLM # containers sharing GPU1 there was no slack left to grow into. # 0.52 leaves ~7.7 GiB free vs ~4.6 GiB at 0.55. Higher needs fewer tenants. # (0.58 failed even earlier — OOM during CUDA graph capture at startup.) MOG_GPU_MEM_UTIL=0.52 MOG_MAX_NUM_SEQS=16 MOG_MAX_NUM_BATCHED_TOKENS=16384 MOG_KV_CACHE_DTYPE=fp8 # ── Vision ────────────────────────────────────────────────────────────────── # 4194304 px = 2048x2048 -> ~5125 image tokens. See the compose comment: the # model's own preprocessor declares 4096x4096, which is both wasteful and fatal # on builds that enforce the image-token count check. MOG_MM_PROCESSOR_KWARGS={"size": {"longest_edge": 4194304, "shortest_edge": 65536}} MOG_LIMIT_MM={"image": 4} # ── Misc ──────────────────────────────────────────────────────────────────── MOG_REASONING_PARSER=qwen3 MOG_REASONING_EFFORT=medium # Intentionally EMPTY — the validated config ran without expandable_segments. MOG_ALLOC_CONF=