feat(mog-sec): quant + serve M.O.G.-SEC pen-test seat; PPL on gen; retire fable
Autonomous overnight run under the operator's full-autonomy grant. End state:
fleet up, gen seat untouched, a new verified pen-test seat serving where fable was.
PPL on the orcarouter gen seat (fable downed to free GPU1 for a nospec probe,
probe torn down after): mean 7.07 / median 5.76, within noise of heresy 6.910 /
5.625 and identical to our recipe's usual 7.059. The gen-seat search is settled.
M.O.G.-SEC: chose Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16 (rev deede677)
over the pre-made ModelOpt NVFP4, which was disqualified on W4A4 4-bit activations
(the AEON degradation mode, catastrophic on a 1M-context model), zero MTP tensors,
and ModelOpt format. Pulled, format-screened (P(<think>) 1.11e-05, clean), quanted
in-house to mixed NVFP4+FP8 (23.4 GB, MTP + vision preserved), and served in the
retired fable slot.
stacks/mog-sec ana-ml2 GPU1 :8019, KV 418,218 tok / 1.60x @ 262K
aliases mog-sec (non-thinking), mog-sec-reasoning (thinking)
gates surface 6/6, MTP 55.3%, format 0/15 leak, vision 7/3/1,
capability 4/4 (delivers offensive-security content)
Served at native 262K, NOT the card's 1M -- the 1M needs YaRN (absent from the
weights' config) plus the SGLang/DFlash2 path the repo ships a deployment kit for,
neither of which is our vLLM surface. A real 1M seat is a separate SGLang project.
Retired char-rp-reasoning + char-rp-fable (zero traffic, pointed at the downed
fable :8019; now 404 cleanly, not repointed -- a security model is not an RP model).
char-rp (meromero) untouched. Vision preprocessor built from the model's own
image_processor block, same trick as the MeroMero seat.
GPU0 seats (gen, meromero) were untouched and healthy throughout. The quant ran in
GPU1 free space with no production seat stopped except fable, which was replaced.
This commit is contained in:
@@ -259,33 +259,52 @@ model_list:
|
||||
# top_p 0.95 / top_k 20) and are unchanged from the DS entry. Verified the FF
|
||||
# chat template honours `enable_thinking` (chat_template.jinja:44) rather than
|
||||
# ignoring it — the mismatch that returned null content on the MeroMero seat.
|
||||
- model_name: char-rp-reasoning
|
||||
# --- char-rp-reasoning + char-rp-fable RETIRED 2026-08-21. Both routed to the
|
||||
# throwaway fablefusion-charrp-probe seat on :8019, which was downed and its
|
||||
# GPU1 slot reassigned to the mog-sec pen-test seat below. Both aliases had
|
||||
# ZERO traffic in the 4-day window before retirement. The RP-reasoning
|
||||
# capability's real home is darkscarlett-charrp-reasoning (:8018, compose
|
||||
# down, weights intact) if it is ever wanted back. Not repointed to mog-sec
|
||||
# -- a security model is not an RP-reasoning model (no false aliases). ---
|
||||
|
||||
# --- mog-sec -> M.O.G.-SEC-27B pen-test seat (ana-ml2 GPU1 :8019, in the retired
|
||||
# fable slot). Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX, stock-Qwen3.8-27B
|
||||
# base, quantized in-house to mixed NVFP4+FP8 with MTP + vision preserved.
|
||||
# Served at native 262K (NOT the card's 1M -- that needs YaRN + SGLang/DFlash2,
|
||||
# not our vLLM path). presence_penalty deliberately 0.0, NOT the fleet's 1.5:
|
||||
# this is a code/security tool and the anti-repetition penalty fights code
|
||||
# structure (and upstream warns it can cause language mixing). Non-thinking. ---
|
||||
- model_name: mog-sec
|
||||
litellm_params:
|
||||
model: hosted_vllm/char-rp-probe
|
||||
model: hosted_vllm/mog-sec-27b
|
||||
api_base: http://10.250.50.54:8019/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
temperature: 1.0
|
||||
top_p: 0.95
|
||||
temperature: 0.7
|
||||
top_p: 0.8
|
||||
presence_penalty: 0.0
|
||||
extra_body:
|
||||
top_k: 20
|
||||
min_p: 0.0
|
||||
repetition_penalty: 1.0
|
||||
chat_template_kwargs:
|
||||
enable_thinking: true
|
||||
enable_thinking: false
|
||||
model_info:
|
||||
mode: chat
|
||||
# char-rp-fable -> the SAME Fable-Fusion seat under its own honest name, so the
|
||||
# evaluation can address it without relying on the temporary repoint above.
|
||||
# Distinct model_name = distinct litellm_params object, which avoids the
|
||||
# shared-deployment param mutation that bleeds sampler overrides between
|
||||
# variants.
|
||||
- model_name: char-rp-fable
|
||||
# mog-sec-reasoning -> the SAME seat, thinking ON. Distinct served-name so a
|
||||
# thinking-off request can't mutate this deployment's enable_thinking (the
|
||||
# shared-config clobber). Canonical Qwen3.8 thinking samplers (temp 1.0/top_p 0.95).
|
||||
- model_name: mog-sec-reasoning
|
||||
litellm_params:
|
||||
model: hosted_vllm/char-rp-probe
|
||||
model: hosted_vllm/mog-sec-27b-thinking
|
||||
api_base: http://10.250.50.54:8019/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
temperature: 1.0
|
||||
top_p: 0.95
|
||||
presence_penalty: 0.0
|
||||
extra_body:
|
||||
top_k: 20
|
||||
min_p: 0.0
|
||||
repetition_penalty: 1.0
|
||||
chat_template_kwargs:
|
||||
enable_thinking: true
|
||||
model_info:
|
||||
|
||||
@@ -0,0 +1,20 @@
|
||||
# mog-sec tunables — pen-test seat (ana-ml2 GPU 1, :8019). Edit the real .env on
|
||||
# the server, never commit it.
|
||||
MOG_IMAGE=vllm/vllm-openai:nightly-311b3513af33bc29b4acb2fde2e9313e5e9966a0
|
||||
API_KEY=
|
||||
MOG_GPU_ID=1
|
||||
|
||||
MOG_CONTAINER_NAME=vllm-mog-sec
|
||||
MOG_PORT=8019
|
||||
MOG_SERVED_NAME=mog-sec-27b
|
||||
MOG_SERVED_NAME_THINK=mog-sec-27b-thinking
|
||||
MOG_MODEL=/tank/aimodels/mog-sec-27b-nvfp4-mixed
|
||||
MOG_QUANT=compressed-tensors
|
||||
MOG_GPU_MEM_UTIL=0.44
|
||||
MOG_MAX_MODEL_LEN=262144
|
||||
MOG_MAX_NUM_SEQS=16
|
||||
MOG_KV_CACHE_DTYPE=fp8
|
||||
MOG_REASONING_PARSER=qwen3
|
||||
MOG_REASONING_EFFORT=medium
|
||||
MOG_SPEC_METHOD=qwen3_5_mtp
|
||||
MOG_SPEC_TOKENS=3
|
||||
@@ -0,0 +1,105 @@
|
||||
# mog-sec — the pen-test seat on ana-ml2 GPU 1 (:8019), in the slot the retired
|
||||
# fablefusion-charrp-probe used to occupy.
|
||||
#
|
||||
# Serves Blackfrost-Research/M.O.G.-SEC-27B-1M-CTX-BF16 (stock-Qwen3.8-27B-based,
|
||||
# vision-intact Qwen3_5ForConditionalGeneration, base-graft MTP head), quantized
|
||||
# in-house to mixed NVFP4 W4A4 (bulk MLP) + FP8 W8A8 (attn / linear_attn / lm_head /
|
||||
# top MLP layers), with the bf16 MTP head grafted back. ⚠ the grafted MTP requires
|
||||
# `re:^mtp.*` in config.json quantization_config.ignore or vLLM loads it uninitialised
|
||||
# (0% accept) — handled by post_quant.py, verified present at build time.
|
||||
#
|
||||
# CONTEXT: served at native 262K, NOT the card's 1M. The 1M needs YaRN rope_scaling
|
||||
# (absent from the weights' config) plus the SGLang/DFlash2 attention path the repo
|
||||
# ships a deployment kit for — neither is our vLLM serving surface. 262K is the honest
|
||||
# native ceiling here; a real 1M seat would be a separate SGLang project.
|
||||
#
|
||||
# Two served-names (base + `-thinking`): LiteLLM keys deployments by (model, api_base),
|
||||
# so mog-sec and mog-sec-reasoning use distinct names to avoid the shared-config
|
||||
# enable_thinking clobber. Same pinned nightly as the gen seat (carries the #51113
|
||||
# qwen3_5_mtp x GDN partial-accept fix that the MTP-on config depends on).
|
||||
|
||||
name: mog-sec
|
||||
|
||||
services:
|
||||
vllm-mog-sec:
|
||||
image: ${MOG_IMAGE:-vllm/vllm-openai:latest}
|
||||
container_name: ${MOG_CONTAINER_NAME:-vllm-mog-sec}
|
||||
restart: unless-stopped
|
||||
ipc: host
|
||||
ports:
|
||||
- "${MOG_PORT:-8019}:8000"
|
||||
volumes:
|
||||
- /tank/aimodels/huggingface:/hfcache
|
||||
- ${MOG_MODEL:-/tank/aimodels/mog-sec-27b-nvfp4-mixed}:/model:ro
|
||||
environment:
|
||||
- HF_HOME=/hfcache
|
||||
- HF_HUB_CACHE=/hfcache/hub
|
||||
- VLLM_API_KEY=${API_KEY:-}
|
||||
- PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
|
||||
command:
|
||||
- /model
|
||||
- --served-model-name
|
||||
- ${MOG_SERVED_NAME:-mog-sec-27b}
|
||||
- ${MOG_SERVED_NAME_THINK:-mog-sec-27b-thinking}
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8000"
|
||||
- --quantization
|
||||
- ${MOG_QUANT:-compressed-tensors}
|
||||
- --gpu-memory-utilization
|
||||
- ${MOG_GPU_MEM_UTIL:-0.44}
|
||||
- --max-model-len
|
||||
- ${MOG_MAX_MODEL_LEN:-262144}
|
||||
- --max-num-seqs
|
||||
- ${MOG_MAX_NUM_SEQS:-16}
|
||||
- --max-num-batched-tokens
|
||||
- "16384"
|
||||
- --trust-remote-code
|
||||
- --dtype
|
||||
- auto
|
||||
- --mamba-cache-dtype
|
||||
- float32
|
||||
- --kv-cache-dtype
|
||||
- ${MOG_KV_CACHE_DTYPE:-fp8}
|
||||
- --enable-prefix-caching
|
||||
- --enable-chunked-prefill
|
||||
- --limit-mm-per-prompt
|
||||
- '{"image": 4}'
|
||||
- --reasoning-parser
|
||||
- ${MOG_REASONING_PARSER:-qwen3}
|
||||
- --default-chat-template-kwargs
|
||||
- '{"reasoning_effort": "${MOG_REASONING_EFFORT:-medium}"}'
|
||||
- --enable-auto-tool-choice
|
||||
- --tool-call-parser
|
||||
- qwen3_coder
|
||||
- --speculative-config
|
||||
- '{"method": "${MOG_SPEC_METHOD:-qwen3_5_mtp}", "num_speculative_tokens": ${MOG_SPEC_TOKENS:-3}}'
|
||||
deploy:
|
||||
resources:
|
||||
reservations:
|
||||
devices:
|
||||
- driver: nvidia
|
||||
device_ids:
|
||||
- "${MOG_GPU_ID:-1}"
|
||||
capabilities:
|
||||
- gpu
|
||||
healthcheck:
|
||||
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 900s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI - Inference
|
||||
- homepage.name=M.O.G.-SEC 27B (pen-test)
|
||||
- homepage.icon=mdi-shield-lock
|
||||
- homepage.description=Uncensored security model, Qwen3.8-27B NVFP4+MTP, 262K — the `mog-sec` seat (ana-ml2 GPU 1)
|
||||
- homepage.href=http://10.250.50.54:${MOG_PORT:-8019}/docs
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
Reference in New Issue
Block a user