Files
vh 019ccff7e8 feat(gemma4-trainee-bench): BF16 bench stack; record that gen's footprint grows with uptime
Adds an ephemeral stack for serving the BF16 trainee base on :8016 under the
char-rp aliases, so the abliterated base can be measured on the same battery and
the same gateway routes as the served seat with no harness edit.

It is a separate stack rather than another variable on gemma4-charrp because
that compose hardcodes `--quantization compressed-tensors` for the NVFP4 build.
Pointing it at unquantized BF16 weights crash-loops immediately —
`TypeError: CompressedTensorsConfig.__init__() missing 3 required positional
arguments: 'target_scheme_map', 'ignore', 'quant_format'` — vLLM trying to read
a quantization config out of a checkpoint that has none. 35 restarts before it
was caught. `restart: "no"` here so a bench seat cannot resurrect itself and
block gen's restore, and no homepage labels so it leaves no permanently-offline
dashboard card.

It cannot coexist with gen and says so: 48.07 GiB of BF16 weights plus gen's
footprint exceeds the 94.97 GiB card before any KV cache. Running it means gen
is stopped.

THE MORE USEFUL FINDING is in the meromero env note: gen's memory footprint
GROWS WITH UPTIME. Measured today at 46,726 MiB (45.6 GiB) after ~3 days up, and
39,424 MiB (38.5 GiB) immediately after a restart — same container, same
--gpu-memory-utilization 0.43, ~7 GiB apart. That is the missing half of this
afternoon's crash-loop: the char-rp seat "fit on the 21st and stopped fitting on
the 24th" because nothing about char-rp changed and gen crept up underneath it.
Headroom arithmetic done against a long-running gen is measuring a moving
number, so the note now says to measure against a freshly-restarted one.

Operator's requested end state reached and verified through the gateway: gen and
summarizer both 200, char-rp down deliberately to hold GPU0 headroom for the
upcoming trainee run, bench seat stopped.
2026-08-24 14:19:49 -07:00
..

meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)

The non-thinking, multimodal RP prose seat. Serves the LiteLLM char-rp alias.

  • Model: G4-MeroMero-v2-31B-NVFP4A16 (Gemma-4-31B, home-quantized weight-only NVFP4A16).
  • Host/GPU: ana-ml2, GPU0 (co-located with gen / vllm-aeon-gen).
  • Port: :8016 → LiteLLM char-rp.
  • Context: 256K (--max-model-len 262144). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context.
  • Vision: enabled (image + text). preprocessor_config.json was materialized from the model's own processor_config.json (Gemma4ImageProcessor); audio is config-declared but weightless.
  • Tool-calling: enabled via the gemma4 parser (not qwen3_coder — that's the Qwen-family XML the other seats use). Gemma-4 emits its own native <|tool_call>call:name{...}<tool_call|> syntax.

Tool-calling — the three flags are a set, don't split them

- --tool-call-parser gemma4        # native <|tool_call> syntax; without it ANY tools request 400s
- --enable-auto-tool-choice
- --reasoning-parser gemma4        # absorbs the <|channel>…<channel|> thought markers
- --default-chat-template-kwargs '{"enable_thinking": false}'   # MANDATORY, see below

Why the last one is mandatory: the gemma4 parser reads enable_thinking out of chat_template_kwargs and defaults it to True (vllm/parser/gemma4.py:439). With True, is_reasoning_end() returns False at a new turn, which pre-initialises the parser engine to REASONING — so all plain RP prose lands in reasoning_content and content comes back null, breaking every char-rp consumer. This model's chat_template.jinja:350 already defaults enable_thinking to false, so passing it explicitly renders a byte-identical prompt (verified across plain / tools / post-tool-response / system-prompt shapes) — it changes nothing about generation, it only corrects the parser's state machine.

Without --reasoning-parser gemma4, the post-tool-response turn leaks a literal <|channel>thought\n<channel|> prefix into content (upstream vllm #45834 — the chat template leaves the prompt sitting inside an open channel block).

Verified green after the fix: tool call (streaming + non-streaming), tool-result round-trip, plain prose in content, vision.

  • Tuning: .envMEROMERO_GPU_MEM_UTIL=0.52 (leaves ~4.6 GB GPU0 headroom), MEROMERO_MAX_MODEL_LEN=262144, MEROMERO_GPU_ID=0.

Replaces the retired char-rp-gguf (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline lives in ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/.

Deploy

scripts/deploy-stack.sh ana-ml2 meromero-charrp     # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d