feat(qwen36-vl): split thinking — non-thinking default + qwen3.6-35b-a3b-thinking variant

The qwen3.6-35b-a3b VL checkpoint is a single hybrid model with a per-
request enable_thinking switch (Qwen3-style), defaulting thinking ON.
Make the default non-thinking and add an opt-in reasoning variant,
mirroring the existing glm-5.1 / glm-5.1-reasoning gateway split.

- qwen36-vl compose: add --reasoning-parser qwen3 (model-matched) so the
  single :8007 endpoint splits <think> into reasoning_content when on and
  routes all output to content when off — serving both modes cleanly.
- litellm gateway: base qwen3.6-35b-a3b pins chat_template_kwargs
  enable_thinking=false (non-thinking default); new qwen3.6-35b-a3b-thinking
  pins enable_thinking=true (opt-in reasoning). Same upstream checkpoint,
  no extra VRAM/container.

Deployed + verified on ana-ml2 (vLLM recreated, healthy) and ana-docker
(litellm reloaded): default returns a direct answer with no reasoning_content;
-thinking returns cleanly-separated reasoning_content, no raw tag leak.
This commit is contained in:
vh
2026-06-15 13:55:11 -07:00
parent 0943d145fb
commit 6de0844323
2 changed files with 45 additions and 1 deletions
+19
View File
@@ -27,6 +27,18 @@
# cost. 0.42 (~40 GB) = weights + graph + generous KV. Granite drops to 0.25/64K
# to make room (the FP8-vs-maxed-granite tradeoff, operator-approved 2026-06-14).
#
# THINKING TOGGLE: this is ONE hybrid checkpoint (not separate Instruct/Thinking
# downloads) with a Qwen3-style per-request `enable_thinking` switch. The chat
# template defaults thinking ON (`<think>\n`); passing
# `chat_template_kwargs={"enable_thinking":false}` emits the empty
# `<think>\n\n</think>\n\n` block (no reasoning). We run --reasoning-parser qwen3
# (model-matched — its vLLM docstring describes THIS checkpoint) so ONE endpoint
# serves BOTH modes cleanly: thinking-ON splits <think>…</think> into
# reasoning_content; thinking-OFF routes everything to content. The gateway
# selects the mode per model_name (stacks/litellm/conf/config.yaml):
# qwen3.6-35b-a3b → enable_thinking:false (non-thinking DEFAULT)
# qwen3.6-35b-a3b-thinking → enable_thinking:true (opt-in reasoning)
#
# All tunables live in .env — edit that, not this file.
name: qwen36-vl
@@ -72,6 +84,13 @@ services:
- --dtype
- auto
- --enable-prefix-caching
# Model-matched reasoning parser for the hybrid thinking toggle (see header).
# Splits <think>…</think> into reasoning_content when thinking is ON; routes
# all output to content when the empty think-block signals thinking OFF — so
# this single :8007 endpoint serves both the non-thinking default and the
# qwen3.6-35b-a3b-thinking gateway variant.
- --reasoning-parser
- qwen3
deploy:
resources:
reservations: