fix(vllm-phi4): Ollama-matching chat template to recover R15 baseline

vLLM's official Phi-4 tokenizer template emits <|end|> after the system turn;
Ollama's does not. That single boundary token regressed brokkr's R15 P02
admission eval (type macro-F1 -33pp) vs the Ollama-measured canonical, while
valid_format held at 1.0. Operator chose to make vLLM match Ollama's leaner
scaffold globally (baseline == production). Adds conf/phi4-chat-template.jinja
(drops the system <|end|>) + mounts it + --chat-template on vllm-phi4. Applied
prompt verified via tokenize/detokenize; brokkr re-smokes probe_vllm.yaml.
This commit is contained in:
vh
2026-06-04 00:29:05 -07:00
parent 40a374b809
commit 90e08f0502
2 changed files with 29 additions and 0 deletions
+5
View File
@@ -210,6 +210,7 @@ services:
- "${PHI4_PORT}:8000"
volumes:
- /tank/aimodels/huggingface:/hfcache
- /opt/docker/conf/vllm/phi4-chat-template.jinja:/config/phi4-chat-template.jinja:ro
environment:
- HF_HOME=/hfcache
- HF_HUB_CACHE=/hfcache/hub
@@ -237,6 +238,10 @@ services:
# FP8 KV cache — halves KV memory at 128K ctx on Ada (cc 8.9); near-lossless.
- --kv-cache-dtype
- ${PHI4_KV_CACHE_DTYPE}
# Ollama-matching scaffold (drops the system-turn <|end|>) so vLLM reproduces
# brokkr's R15 canonical baseline. See conf/phi4-chat-template.jinja.
- --chat-template
- /config/phi4-chat-template.jinja
deploy:
resources:
reservations: