Operator approved after real-use testing. The experimental standalone container is retired and stacks/mog-sec is canonical again, with restart: unless-stopped so the configuration survives a reboot. Cutover verified against the container it replaces: KV pool 526,617 tokens at 1.10x concurrency, identical; zero restarts; both gateway aliases serving; DFlash2 confirmed drafting at k=7 with 231 draft tokens over 33 drafts; vision working at 2048x2048. One variable was deliberately dropped rather than carried over. The previous stack hardcoded PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, the validated container never set it, and the quant playbook records expandable_segments corrupting retained tensors in another context. The compose now defaults it empty via MOG_ALLOC_CONF. Promoting the stack as it stood would have shipped a variable the tested configuration did not have. The speculative config moves into a single MOG_SPEC_CONFIG carrying the whole JSON, because the two shapes are not interchangeable: dflash requires a model pointing at the drafter and MTP must not have one, so a method-plus-tokens template cannot express both. Also parameterised: MOG_DRAFT_MODEL, MOG_MM_PROCESSOR_KWARGS, MOG_MAX_NUM_BATCHED_TOKENS. The mm-processor image cap is now mandatory rather than incidental. The model's own preprocessor declares 4096x4096, which expands to 16384 image tokens and kills startup on builds that enforce the image-token count check. Adds the .env.example this stack never had, carrying the measured rationale for each value and the one-line rollback.
67 lines
4.1 KiB
Bash
67 lines
4.1 KiB
Bash
# mog-sec — pen-test seat (ana-ml2 GPU1, :8019). Copy to .env on the host.
|
|
#
|
|
# Values below are the configuration VALIDATED 2026-08-22: DFlash2 speculative
|
|
# decoding on a newer vLLM, 480K context, 2048x2048 vision. Promoted from a
|
|
# standalone experimental container after real-use testing.
|
|
|
|
# ── Image ───────────────────────────────────────────────────────────────────
|
|
# Contains DFlash2 (#52816) AND GDN spec-decode fix #53077. Verified by
|
|
# ancestry, not by version string: this reports 0.26.1rc1.dev1102, which LOOKS
|
|
# older than the 0.27.2rc1.dev150 it replaced — that is a setuptools_scm
|
|
# tag-reachability artifact, not a downgrade. It is +259 commits, behind_by=0.
|
|
# ⚠ Bumping this is a CONFIG-COMPATIBILITY event, not just a version change.
|
|
# Newer transformers enforces model-config constraints older ones ignored.
|
|
MOG_IMAGE=vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013
|
|
|
|
API_KEY=
|
|
MOG_GPU_ID=1
|
|
MOG_CONTAINER_NAME=vllm-mog-sec
|
|
MOG_PORT=8019
|
|
MOG_SERVED_NAME=mog-sec-27b
|
|
MOG_SERVED_NAME_THINK=mog-sec-27b-thinking
|
|
MOG_MODEL=/tank/aimodels/mog-sec-27b-nvfp4-mixed
|
|
MOG_QUANT=compressed-tensors
|
|
|
|
# ── Speculative decoding — DFlash2 ──────────────────────────────────────────
|
|
# 2B block-diffusion drafter, drafts a whole 8-token block in one pass. Unlike
|
|
# the MTP head it reads the TARGET's live hidden states (layers 5/19/33/47/61),
|
|
# so it is model-agnostic — the same weights measured identically against a
|
|
# different finetune (3.254 vs 3.252 accepted tok/forward, a 0.06% delta).
|
|
# Measured on this seat: 2.676 -> 3.252 tok/forward, 110.5 -> 130.0 tok/s.
|
|
# ⚠ The drafter is COUPLED to its target and lives in that engine's process:
|
|
# the weights file is shareable across seats, the 3.85 GB of VRAM is NOT.
|
|
MOG_DRAFT_MODEL=/tank/aimodels/qwen38-27b-dflash2-drafter
|
|
MOG_SPEC_CONFIG={"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7}
|
|
# ROLLBACK to the previous behaviour — one line, plus MOG_IMAGE above:
|
|
# MOG_SPEC_CONFIG={"method": "qwen3_5_mtp", "num_speculative_tokens": 3}
|
|
|
|
# ── Context and memory ──────────────────────────────────────────────────────
|
|
# The model carries a complete YaRN config (rope_type yarn, factor 4.0,
|
|
# original_max_position_embeddings 262144, max_position_embeddings 1000000),
|
|
# so context is a KV-MEMORY choice, not a model limit.
|
|
# ⚠ max-model-len must stay UNDER the KV pool or vLLM refuses to start.
|
|
# At 0.55 the pool is ~526,617 tokens -> 480,000 gives 1.10x concurrency.
|
|
# 262144 instead would give ~1.86x. Straight trade: context vs concurrency.
|
|
MOG_MAX_MODEL_LEN=480000
|
|
# ⚠ 0.55 IS THE STABLE CEILING while GPU1's other tenants are up. 0.58 sized a
|
|
# bigger pool and then OOM'd during CUDA graph capture (process reached
|
|
# 57.49 GiB against ~57.6 free). Real 1M context needs ~49 GiB of KV and so
|
|
# requires evicting most of GPU1 — a fleet decision, not a flag.
|
|
MOG_GPU_MEM_UTIL=0.55
|
|
MOG_MAX_NUM_SEQS=16
|
|
MOG_MAX_NUM_BATCHED_TOKENS=16384
|
|
MOG_KV_CACHE_DTYPE=fp8
|
|
|
|
# ── Vision ──────────────────────────────────────────────────────────────────
|
|
# 4194304 px = 2048x2048 -> ~5125 image tokens. See the compose comment: the
|
|
# model's own preprocessor declares 4096x4096, which is both wasteful and fatal
|
|
# on builds that enforce the image-token count check.
|
|
MOG_MM_PROCESSOR_KWARGS={"size": {"longest_edge": 4194304, "shortest_edge": 65536}}
|
|
MOG_LIMIT_MM={"image": 4}
|
|
|
|
# ── Misc ────────────────────────────────────────────────────────────────────
|
|
MOG_REASONING_PARSER=qwen3
|
|
MOG_REASONING_EFFORT=medium
|
|
# Intentionally EMPTY — the validated config ran without expandable_segments.
|
|
MOG_ALLOC_CONF=
|