fix(erp-seat): tool_choice=none returned an empty turn — add --exclude-tools-when-tool-choice-none (vLLM kept the tools in the prompt, the model called one, parsing was off); 12-shape tool matrix green before/after

This commit is contained in:
vh
2026-09-08 23:01:26 -07:00
parent 32399d0da2
commit 7f6be8a56a
3 changed files with 18 additions and 0 deletions
+6
View File
@@ -187,6 +187,12 @@ COMPLETE and gated RESCUED (02:13 PDT). Live open items:_
60 router + 191 vision) from `/tank/aimodels/erp-tune-v6-bf16` (merged-run06 relayed gx10→nh3-dev→ana-ml2 in 17 min, 60 router + 191 vision) from `/tank/aimodels/erp-tune-v6-bf16` (merged-run06 relayed gx10→nh3-dev→ana-ml2 in 17 min,
no key path gx10→ana-ml2). Quant = `services/erp-seat-quant/` — DATA-FREE (~90 s; playbook §3.16), tokenizer cap no key path gx10→ana-ml2). Quant = `services/erp-seat-quant/` — DATA-FREE (~90 s; playbook §3.16), tokenizer cap
reset, encode-equal to source. Smoke: prose in `content`, ~200 tok/s single-stream (n=3, spread <1%). reset, encode-equal to source. Smoke: prose in `content`, ~200 tok/s single-stream (n=3, spread <1%).
**Tool calling FIXED 2026-09-08 23:00 PT (operator: "fix toolcalling with the trial seat"):** the one reproducible
defect in a 12-shape matrix was `tool_choice:"none"` → empty turn (content AND tool_calls null, 3/3) — vLLM kept the
tools in the prompt, the model called one, parsing was off. Fix = `--exclude-tools-when-tool-choice-none` (seat
recreated; none→prose 3/3; auto/required/named/parallel/nested/empty-list/streaming/round-trip all green).
⚠ `stacks/gemma4-charrp` (char-rp seat) has the same exposure and no flag yet — bouncing it is a consumer-visible
restart, so left for the operator's word.
⚠ NOT gate-parity (gated artifact = bf16 arm on the GX10; this seat has a smoke test + 3-run decode probe only) — ⚠ NOT gate-parity (gated artifact = bf16 arm on the GX10; this seat has a smoke test + 3-run decode probe only) —
**operator ruled "no gate"**; the config block states it as unrated on every safety axis. **operator ruled "no gate"**; the config block states it as unrated on every safety axis.
⚠ **ana-ml2 mesh return routes are NON-PERSISTENT** (`ip route replace 10.100/16, 10.0/16, 10.6.110/24, 100.64/10 ⚠ **ana-ml2 mesh return routes are NON-PERSISTENT** (`ip route replace 10.100/16, 10.0/16, 10.6.110/24, 100.64/10
+7
View File
@@ -12,6 +12,13 @@ checkpoint so the GX10 is free to train the next run. First occupant: **run 6**
and reasoning parsers, `enable_thinking` pinned false, the model's own stock template and reasoning parsers, `enable_thinking` pinned false, the model's own stock template
(`ae53464b…`, the one it trained through). Without the reasoning parser the post-tool turn leaks (`ae53464b…`, the one it trained through). Without the reasoning parser the post-tool turn leaks
`<|channel>` markers; without the kwargs pin all prose lands in `reasoning_content`. `<|channel>` markers; without the kwargs pin all prose lands in `reasoning_content`.
- **`tool_choice: "none"` trap (measured 2026-09-08, fixed with `--exclude-tools-when-tool-choice-none`).**
Without the flag vLLM still renders the tools into the prompt, the model emits a tool call
anyway, and because parsing is off for `none` the reply is `content: null, tool_calls: null`
— an empty turn, 3/3 reproductions. With the flag the tools are dropped from the prompt and
the model answers in prose (3/3). The rest of the matrix (auto / required / named / parallel /
nested schema / empty `tools: []` / streaming / tool-result round trip) was green before and
after. `stacks/gemma4-charrp` has the same exposure and does NOT carry the flag yet.
- **GPU1 is shared** — check real usage (`nvidia-smi --query-compute-apps=pid,used_memory`) before - **GPU1 is shared** — check real usage (`nvidia-smi --query-compute-apps=pid,used_memory`) before
raising `ERP_GPU_MEM_UTIL`; the flag sizes KV, not CUDA context. raising `ERP_GPU_MEM_UTIL`; the flag sizes KV, not CUDA context.
- **Rollback / next run:** point `ERP_MODEL` + `ERP_SERVED_NAME` at the next quant dir, keep the - **Rollback / next run:** point `ERP_MODEL` + `ERP_SERVED_NAME` at the next quant dir, keep the
+5
View File
@@ -35,6 +35,11 @@ services:
- gemma4 - gemma4
- --default-chat-template-kwargs - --default-chat-template-kwargs
- '{"enable_thinking": false}' - '{"enable_thinking": false}'
# tool_choice:"none" trap (measured 2026-09-08): without this flag vLLM still renders the
# tools into the prompt, the model emits a tool call anyway, and because tool parsing is
# off for tool_choice=none the reply comes back with content=null AND tool_calls=null —
# an empty turn. The flag drops the tools from the prompt so the model answers in prose.
- --exclude-tools-when-tool-choice-none
- --chat-template - --chat-template
- ${ERP_CHAT_TEMPLATE:-/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja} - ${ERP_CHAT_TEMPLATE:-/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja}
- --max-model-len - --max-model-len