feat(mog-sec): promote the DFlash2 configuration into the compose stack

Operator approved after real-use testing. The experimental standalone
container is retired and stacks/mog-sec is canonical again, with
restart: unless-stopped so the configuration survives a reboot.

Cutover verified against the container it replaces: KV pool 526,617 tokens
at 1.10x concurrency, identical; zero restarts; both gateway aliases
serving; DFlash2 confirmed drafting at k=7 with 231 draft tokens over 33
drafts; vision working at 2048x2048.

One variable was deliberately dropped rather than carried over. The previous
stack hardcoded PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, the
validated container never set it, and the quant playbook records
expandable_segments corrupting retained tensors in another context. The
compose now defaults it empty via MOG_ALLOC_CONF. Promoting the stack as it
stood would have shipped a variable the tested configuration did not have.

The speculative config moves into a single MOG_SPEC_CONFIG carrying the
whole JSON, because the two shapes are not interchangeable: dflash requires
a model pointing at the drafter and MTP must not have one, so a
method-plus-tokens template cannot express both. Also parameterised:
MOG_DRAFT_MODEL, MOG_MM_PROCESSOR_KWARGS, MOG_MAX_NUM_BATCHED_TOKENS.

The mm-processor image cap is now mandatory rather than incidental. The
model's own preprocessor declares 4096x4096, which expands to 16384 image
tokens and kills startup on builds that enforce the image-token count check.

Adds the .env.example this stack never had, carrying the measured rationale
for each value and the one-line rollback.
This commit is contained in:
vh
2026-08-22 01:16:27 -07:00
parent 20ac53052b
commit 8389470898
4 changed files with 108 additions and 21 deletions
+31 -4
View File
@@ -31,11 +31,20 @@ services:
volumes:
- /tank/aimodels/huggingface:/hfcache
- ${MOG_MODEL:-/tank/aimodels/mog-sec-27b-nvfp4-mixed}:/model:ro
# DFlash2 speculative drafter. Mounted unconditionally — it is inert if
# MOG_SPEC_CONFIG selects an MTP method that does not reference /drafter.
- ${MOG_DRAFT_MODEL:-/tank/aimodels/qwen38-27b-dflash2-drafter}:/drafter:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
- VLLM_API_KEY=${API_KEY:-}
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
# ⚠ DEFAULTS TO UNSET, deliberately. This seat previously hardcoded
# PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. The DFlash2 config
# validated 2026-08-22 ran WITHOUT it, and the quant playbook §3.10
# records expandable_segments corrupting retained tensors in another
# context. Do not re-enable it casually — that would ship a variable the
# tested configuration did not have.
- PYTORCH_CUDA_ALLOC_CONF=${MOG_ALLOC_CONF:-}
command:
- /model
- --served-model-name
@@ -54,7 +63,9 @@ services:
- --max-num-seqs
- ${MOG_MAX_NUM_SEQS:-16}
- --max-num-batched-tokens
- "16384"
# ⚠ Raising this costs peak-activation VRAM straight out of the KV pool
# (measured 2026-08-22: 16384 -> 32768 cost ~3 GiB of KV for no benefit).
- ${MOG_MAX_NUM_BATCHED_TOKENS:-16384}
- --trust-remote-code
- --dtype
- auto
@@ -65,7 +76,15 @@ services:
- --enable-prefix-caching
- --enable-chunked-prefill
- --limit-mm-per-prompt
- '{"image": 4}'
- '${MOG_LIMIT_MM:-{"image": 4}}'
# ⚠ MANDATORY on a newer vLLM. The model's own preprocessor_config.json
# declares size.longest_edge = 16777216 px (4096x4096), which expands to
# 16384 image tokens — one image eating 6% of a 262K context, and enough
# to kill startup on builds that enforce the text-vs-ids count check.
# This caps the dummy profiling image AND real images. Cost scales as
# (edge/patch)^2 / merge^2, so 2048x2048 -> ~5125 tokens.
- --mm-processor-kwargs
- '${MOG_MM_PROCESSOR_KWARGS:-{"size": {"longest_edge": 4194304, "shortest_edge": 65536}}}'
- --reasoning-parser
- ${MOG_REASONING_PARSER:-qwen3}
- --default-chat-template-kwargs
@@ -73,8 +92,16 @@ services:
- --enable-auto-tool-choice
- --tool-call-parser
- qwen3_coder
# ONE env var carrying the whole JSON, because the two speculative shapes
# are not interchangeable: dflash needs a "model" pointing at the drafter,
# MTP must NOT have one. A method+tokens template cannot express both.
# DFlash2 : {"method": "dflash", "model": "/drafter", "num_speculative_tokens": 7}
# MTP : {"method": "qwen3_5_mtp", "num_speculative_tokens": 3}
# ⚠ Comparing the two requires matching num_speculative_tokens — see the
# quant playbook §5.1: MTP runs a single-module head autoregressively, so
# deeper k improves acceptance and DESTROYS throughput.
- --speculative-config
- '{"method": "${MOG_SPEC_METHOD:-qwen3_5_mtp}", "num_speculative_tokens": ${MOG_SPEC_TOKENS:-3}}'
- '${MOG_SPEC_CONFIG:-{"method": "qwen3_5_mtp", "num_speculative_tokens": 3}}'
deploy:
resources:
reservations: