# erp-seat — ERP-tune seat on ana-ml2 (GPU1, `:8021`) Serves the latest gated ERP LoRA merge as an **NVFP4A16** (weight-only) compressed-tensors checkpoint so the GX10 is free to train the next run. First occupant: **run 6** — `erp-tune-v6-nvfp4a16` = merged-run06 (jenerallee78 ARA-abliterated Gemma-4-26B-A4B-it, index `33c59654…`, + R47 SFT r6) quantized by `services/erp-seat-quant/`. - **True name only.** `--served-model-name erp-tune-v6-nvfp4a16`. Gateway aliases (`trial`) are set in LiteLLM on the operator's word, never here (no silent substitution — the bf16 arm on the GX10 and this NVFP4 arm are different artifacts). - **Recipe** = `stacks/gemma4-charrp` (same arch + format, proven on this box): `gemma4` tool and reasoning parsers, `enable_thinking` pinned false, the model's own stock template (`ae53464b…`, the one it trained through). Without the reasoning parser the post-tool turn leaks `<|channel>` markers; without the kwargs pin all prose lands in `reasoning_content`. - **`tool_choice: "none"` trap (measured 2026-09-08, fixed with `--exclude-tools-when-tool-choice-none`).** Without the flag vLLM still renders the tools into the prompt, the model emits a tool call anyway, and because parsing is off for `none` the reply is `content: null, tool_calls: null` — an empty turn, 3/3 reproductions. With the flag the tools are dropped from the prompt and the model answers in prose (3/3). The rest of the matrix (auto / required / named / parallel / nested schema / empty `tools: []` / streaming / tool-result round trip) was green before and after. `stacks/gemma4-charrp` has the same exposure and does NOT carry the flag yet. - **GPU1 is shared** — check real usage (`nvidia-smi --query-compute-apps=pid,used_memory`) before raising `ERP_GPU_MEM_UTIL`; the flag sizes KV, not CUDA context. - **Rollback / next run:** point `ERP_MODEL` + `ERP_SERVED_NAME` at the next quant dir, keep the previous on disk. Deploy with `scripts/deploy-stack.sh ana-ml2 erp-seat`.