Files
esh-pfi-infrastructure/stacks/vllm/conf/phi4-chat-template.jinja
T
vh 90e08f0502 fix(vllm-phi4): Ollama-matching chat template to recover R15 baseline
vLLM's official Phi-4 tokenizer template emits <|end|> after the system turn;
Ollama's does not. That single boundary token regressed brokkr's R15 P02
admission eval (type macro-F1 -33pp) vs the Ollama-measured canonical, while
valid_format held at 1.0. Operator chose to make vLLM match Ollama's leaner
scaffold globally (baseline == production). Adds conf/phi4-chat-template.jinja
(drops the system <|end|>) + mounts it + --chat-template on vllm-phi4. Applied
prompt verified via tokenize/detokenize; brokkr re-smokes probe_vllm.yaml.
2026-06-04 00:29:05 -07:00

25 lines
999 B
Django/Jinja

{#-
Phi-4-mini chat template — OLLAMA-MATCHING variant.
Why this exists: the official HF tokenizer template emits <|end|> after the
SYSTEM turn; Ollama's phi4 template does NOT (its system-block <|end|> is only
in the tools branch). That single boundary token regressed brokkr's R15 P02
admission eval on vLLM vs the Ollama-measured canonical (type macro-F1 -33pp)
while valid_format held at 1.0. Operator chose (2026-06-04) to make vLLM match
Ollama's leaner scaffold globally so baseline == production.
Renders (system + user, add_generation_prompt):
<|system|>{sys}<|user|>{usr}<|end|><|assistant|>
i.e. NO <|end|> after the system turn (the only delta from official).
-#}
{%- for message in messages -%}
{%- if message['role'] == 'system' -%}
{{- '<|system|>' + message['content'] -}}
{%- else -%}
{{- '<|' + message['role'] + '|>' + message['content'] + '<|end|>' -}}
{%- endif -%}
{%- endfor -%}
{%- if add_generation_prompt -%}
{{- '<|assistant|>' -}}
{%- endif -%}