feat(mog-sec): promote the DFlash2 configuration into the compose stack

Operator approved after real-use testing. The experimental standalone
container is retired and stacks/mog-sec is canonical again, with
restart: unless-stopped so the configuration survives a reboot.

Cutover verified against the container it replaces: KV pool 526,617 tokens
at 1.10x concurrency, identical; zero restarts; both gateway aliases
serving; DFlash2 confirmed drafting at k=7 with 231 draft tokens over 33
drafts; vision working at 2048x2048.

One variable was deliberately dropped rather than carried over. The previous
stack hardcoded PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True, the
validated container never set it, and the quant playbook records
expandable_segments corrupting retained tensors in another context. The
compose now defaults it empty via MOG_ALLOC_CONF. Promoting the stack as it
stood would have shipped a variable the tested configuration did not have.

The speculative config moves into a single MOG_SPEC_CONFIG carrying the
whole JSON, because the two shapes are not interchangeable: dflash requires
a model pointing at the drafter and MTP must not have one, so a
method-plus-tokens template cannot express both. Also parameterised:
MOG_DRAFT_MODEL, MOG_MM_PROCESSOR_KWARGS, MOG_MAX_NUM_BATCHED_TOKENS.

The mm-processor image cap is now mandatory rather than incidental. The
model's own preprocessor declares 4096x4096, which expands to 16384 image
tokens and kills startup on builds that enforce the image-token count check.

Adds the .env.example this stack never had, carrying the measured rationale
for each value and the one-line rollback.
This commit is contained in:
vh
2026-08-22 01:16:27 -07:00
parent 20ac53052b
commit 8389470898
4 changed files with 108 additions and 21 deletions
@@ -154,11 +154,28 @@ IndexError, workaround is disabling one).
`rope_type: yarn`, `factor: 4.0`, `original_max_position_embeddings: 262144`,
`max_position_embeddings: 1000000`. Context is a KV-memory choice, not a model limit.
## Live state — sec is NOT running from its compose stack
## Live state — PROMOTED to the compose stack 2026-08-22
`vllm-sec-dflash2`, a **standalone container** on sec's port with sec's served names, so the
`sec` / `sec-reasoning` gateway aliases work unchanged. `/opt/docker/compose/mog-sec` is
**stopped but unmodified**.
**Operator-approved after real-use testing** ("performing very well"). The experimental
standalone container is gone; `stacks/mog-sec/` is canonical and `restart: unless-stopped` means
it survives reboots. Cutover verified: **KV pool 526,617 / 1.10x — identical to the container it
replaced**, restarts 0, both gateway aliases serving, DFlash2 confirmed drafting at k=7
(231 draft tokens over 33 drafts), vision working.
⚠ **One variable was deliberately REMOVED, not carried over.** The old stack hardcoded
`PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True`; the validated DFlash2 container never set it,
and playbook §3.10 records expandable_segments corrupting retained tensors elsewhere. The compose
now defaults it EMPTY (`MOG_ALLOC_CONF`). Promoting it as-was would have shipped a variable the
tested configuration did not have.
**Compose is now parameterised for the shapes that differ:** `MOG_SPEC_CONFIG` carries the whole
speculative JSON (dflash needs `"model": "/drafter"`, MTP must not have one — a method+tokens
template cannot express both), plus `MOG_MM_PROCESSOR_KWARGS`, `MOG_DRAFT_MODEL`,
`MOG_MAX_NUM_BATCHED_TOKENS`, `MOG_ALLOC_CONF`.
**ROLLBACK:** `.env.bak-pre-dflash2-20260822` and `compose.yaml.bak-pre-dflash2-20260822` on the
host; or one line — `MOG_SPEC_CONFIG={"method": "qwen3_5_mtp", "num_speculative_tokens": 3}` plus
the old `MOG_IMAGE`.
| | production sec | current |
|---|---|---|
@@ -168,11 +185,8 @@ IndexError, workaround is disabling one).
| KV pool | 418,218 (1.60×) | **526,617 (1.10×)** |
| images | 4096² → 16,384 tok | **2048² → ~5,125 tok** (`--mm-processor-kwargs` size cap) |
**ROLLBACK is two commands:** `docker rm -f vllm-sec-dflash2` then `docker compose up -d` in
`/opt/docker/compose/mog-sec`.
⚠ **`--gpu-memory-utilization 0.55` is the stable ceiling** while GPU1's other tenants are up.
0.58 sized KV at 594,172 then **OOM'd during CUDA graph capture** — the process reached 57.49 GiB
against ~57.6 free. Real 1M context needs ~49 GiB of KV and therefore evicting most of GPU1.
Launcher: `/tmp/run_sec_dflash2.sh` on ana-ml2 (ephemeral — re-derive from this table if lost).
Canonical config: `stacks/mog-sec/{compose.yaml,.env.example}` in this repo.