Full seat rebalance across GPU0/GPU1 (flash on GPU2 and empty GPU3 untouched), operator-directed. Every target seat now serves native 262,144 context with concurrency in the requested 1.2-2.5x band, verified from live boot logs: cyberprev (sec) 262144 @ 1.37x depth-probed CLEAN to 259,722 tokens flash-next (gen) 262144 @ 1.31x (untouched, already in band) gen-small (NEW) 262144 @ 2.56x MTP k=3 measured 69.6% accept / 3.09 len char-rp 262144 @ 1.22x (was 1.02x; util 0.40->0.52) char-rp-fast 262144 @ 2.04x (util cap 0.30->0.24, pinned KV unchanged) - gen-small: NEW seat, Qwen3.6-35B-A3B (3B active MoE), llmfan46 Heretic (MPOA) NVFP4 experts-only, already on disk at qwen36-35b-a3b-heretic-nvfp4. There is no general Qwen3.8 A3B (3.8 MoEs are Flash-Next and the 2.4T), so this is the 3.6 fallback the operator specified. GPU0, :8026, MTP k=3, coherent and MTP-verified before wiring. gen-small / gen-small-reasoning gateway aliases. - coder: 8192 @ 13.75x -> 16384 @ 4.70x (util 0.06->0.055). Context doubled, waste cut. Not the exact 2-3x target: the 1.5B weight+overhead floor (~4.2 GiB) sits just under the util knob's resolution, so hitting <=3x reliably needs a --kv-cache-memory byte pin (compose change) rather than the util fraction. - cyberprev raised 163840 -> 262144: depth-probed with non-repeating prompts to 259,722 tokens, clean (no OOM, memory flat). Unlike mog-sec (same base arch, capped at 163840 for depth crashes), this checkpoint holds native depth. - Gateway (operator calls): summarizer + classifier -> gen-small; new classifier-large -> gen-large (flash) for the accuracy tier; summarizer-large stays on flash. All verified end-to-end. - GPU1 hit its ceiling raising char-rp; resolved by trimming char-rp-fast's reservation cap (its KV is pinned, so concurrency held at 2.04x) rather than moving a utility seat -- the shared GPU_ID on reward/embed/rerank made a single-seat move messier than the in-GPU rebalance. Seat inventory regenerated from the live containers.
39 lines
1.6 KiB
Bash
39 lines
1.6 KiB
Bash
# gen-small-seat — Qwen3.6-35B-A3B Heretic (llmfan46), fv-ml1 GPU 0, :8026.
|
|
# Copy to .env on the host at /opt/docker/compose/gen-small-seat/.env.
|
|
|
|
GEN_SMALL_IMAGE=vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013
|
|
API_KEY=
|
|
|
|
# ── Placement ── shares GPU 0 with cyberprev (sec). KV pinned in bytes so the two
|
|
# co-resident seats don't fight over a util fraction.
|
|
GEN_SMALL_GPU_ID=0
|
|
GEN_SMALL_PORT=8026
|
|
GEN_SMALL_CONTAINER_NAME=vllm-gen-small
|
|
|
|
# ── Model ── llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-NVFP4-Experts-Only
|
|
# Heretic v1.3.0 (MPOA); modelopt NVFP4 experts-only; 19 MTP preserved.
|
|
GEN_SMALL_MODEL=/tank/aimodels/qwen36-35b-a3b-heretic-nvfp4
|
|
GEN_SMALL_QUANT=modelopt_fp4
|
|
GEN_SMALL_SERVED_NAME=gen-small
|
|
GEN_SMALL_SERVED_NAME_THINK=gen-small-thinking
|
|
|
|
# ── Memory ── FIRST-BOOT values, expected to be tuned to a sane concurrency band
|
|
# (1.2-2.5x) by reading the boot log. A3B attention + fp8 KV is cheap, so 8 GiB likely
|
|
# over-provisions; pin down after the first boot reports its token count.
|
|
GEN_SMALL_GPU_MEM_UTIL=0.55
|
|
GEN_SMALL_KV_CACHE_MEMORY=8589934592
|
|
GEN_SMALL_MAX_MODEL_LEN=262144
|
|
GEN_SMALL_MAX_NUM_SEQS=16
|
|
GEN_SMALL_MAX_NUM_BATCHED_TOKENS=4096
|
|
GEN_SMALL_KV_CACHE_DTYPE=fp8
|
|
|
|
# ── Speculative decoding ── MTP k=3. ⚠ Verify acceptance on THIS build (abliteration can
|
|
# desync MTP; MoE MTP has its own failure modes). Delete both --speculative-config lines
|
|
# in the compose to disable.
|
|
GEN_SMALL_SPEC_CONFIG={"method": "qwen3_5_mtp", "num_speculative_tokens": 3}
|
|
|
|
GEN_SMALL_REASONING_PARSER=qwen3
|
|
GEN_SMALL_REASONING_EFFORT=medium
|
|
GEN_SMALL_TOOL_PARSER=qwen3_xml
|
|
GEN_SMALL_ALLOC_CONF=
|