# mog-sec — pen-test seat (ana-ml2 GPU1, :8019). Copy to .env on the host. # # Values below are the configuration VALIDATED 2026-08-22: DFlash2 speculative # decoding on a newer vLLM, 480K context, 2048x2048 vision. Promoted from a # standalone experimental container after real-use testing. # ── Image ─────────────────────────────────────────────────────────────────── # Contains DFlash2 (#52816) AND GDN spec-decode fix #53077. Verified by # ancestry, not by version string: this reports 0.26.1rc1.dev1102, which LOOKS # older than the 0.27.2rc1.dev150 it replaced — that is a setuptools_scm # tag-reachability artifact, not a downgrade. It is +259 commits, behind_by=0. # ⚠ Bumping this is a CONFIG-COMPATIBILITY event, not just a version change. # Newer transformers enforces model-config constraints older ones ignored. MOG_IMAGE=vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013 API_KEY= MOG_GPU_ID=1 MOG_CONTAINER_NAME=vllm-mog-sec MOG_PORT=8019 MOG_SERVED_NAME=mog-sec-27b MOG_SERVED_NAME_THINK=mog-sec-27b-thinking MOG_MODEL=/tank/aimodels/mog-sec-27b-nvfp4-mixed MOG_QUANT=compressed-tensors # ── Speculative decoding — DFlash2 ────────────────────────────────────────── # 2B block-diffusion drafter, drafts a whole 8-token block in one pass. Unlike # the MTP head it reads the TARGET's live hidden states (layers 5/19/33/47/61), # so it is model-agnostic — the same weights measured identically against a # different finetune (3.254 vs 3.252 accepted tok/forward, a 0.06% delta). # Measured on this seat: 2.676 -> 3.252 tok/forward, 110.5 -> 130.0 tok/s. # ⚠ The drafter is COUPLED to its target and lives in that engine's process: # the weights file is shareable across seats, the 3.85 GB of VRAM is NOT. MOG_DRAFT_MODEL=/tank/aimodels/qwen38-27b-dflash2-drafter MOG_SPEC_CONFIG={"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7} # ROLLBACK to the previous behaviour — one line, plus MOG_IMAGE above: # MOG_SPEC_CONFIG={"method": "qwen3_5_mtp", "num_speculative_tokens": 3} # ── Context and memory ────────────────────────────────────────────────────── # The model carries a complete YaRN config (rope_type yarn, factor 4.0, # original_max_position_embeddings 262144, max_position_embeddings 1000000), # so context is a KV-MEMORY choice, not a model limit. # ⚠ max-model-len must stay UNDER the KV pool or vLLM refuses to start. # At 0.55 the pool is ~526,617 tokens -> 480,000 gives 1.10x concurrency. # 262144 instead would give ~1.86x. Straight trade: context vs concurrency. MOG_MAX_MODEL_LEN=480000 # ⚠ 0.55 IS THE STABLE CEILING while GPU1's other tenants are up. 0.58 sized a # bigger pool and then OOM'd during CUDA graph capture (process reached # 57.49 GiB against ~57.6 free). Real 1M context needs ~49 GiB of KV and so # requires evicting most of GPU1 — a fleet decision, not a flag. MOG_GPU_MEM_UTIL=0.55 MOG_MAX_NUM_SEQS=16 MOG_MAX_NUM_BATCHED_TOKENS=16384 MOG_KV_CACHE_DTYPE=fp8 # ── Vision ────────────────────────────────────────────────────────────────── # 4194304 px = 2048x2048 -> ~5125 image tokens. See the compose comment: the # model's own preprocessor declares 4096x4096, which is both wasteful and fatal # on builds that enforce the image-token count check. MOG_MM_PROCESSOR_KWARGS={"size": {"longest_edge": 4194304, "shortest_edge": 65536}} MOG_LIMIT_MM={"image": 4} # ── Misc ──────────────────────────────────────────────────────────────────── MOG_REASONING_PARSER=qwen3 MOG_REASONING_EFFORT=medium # Intentionally EMPTY — the validated config ran without expandable_segments. MOG_ALLOC_CONF=