revert(vllm-phi4): back to canonical/official Phi-4 chat template
Operator chose option (ii): keep the OFFICIAL Phi-4 format globally rather than
impose Ollama's leaner scaffold on every phi4 consumer. Removes the
--chat-template override + the conf/phi4-chat-template.jinja file (90e08f0).
vLLM now uses the tokenizer's built-in template (system <|end|> present);
verified 7-token render via tokenize/detokenize. brokkr re-baselines its R15
canonical on the official scaffold so baseline == production.
This commit is contained in:
@@ -210,7 +210,6 @@ services:
|
|||||||
- "${PHI4_PORT}:8000"
|
- "${PHI4_PORT}:8000"
|
||||||
volumes:
|
volumes:
|
||||||
- /tank/aimodels/huggingface:/hfcache
|
- /tank/aimodels/huggingface:/hfcache
|
||||||
- /opt/docker/conf/vllm/phi4-chat-template.jinja:/config/phi4-chat-template.jinja:ro
|
|
||||||
environment:
|
environment:
|
||||||
- HF_HOME=/hfcache
|
- HF_HOME=/hfcache
|
||||||
- HF_HUB_CACHE=/hfcache/hub
|
- HF_HUB_CACHE=/hfcache/hub
|
||||||
@@ -238,10 +237,6 @@ services:
|
|||||||
# FP8 KV cache — halves KV memory at 128K ctx on Ada (cc 8.9); near-lossless.
|
# FP8 KV cache — halves KV memory at 128K ctx on Ada (cc 8.9); near-lossless.
|
||||||
- --kv-cache-dtype
|
- --kv-cache-dtype
|
||||||
- ${PHI4_KV_CACHE_DTYPE}
|
- ${PHI4_KV_CACHE_DTYPE}
|
||||||
# Ollama-matching scaffold (drops the system-turn <|end|>) so vLLM reproduces
|
|
||||||
# brokkr's R15 canonical baseline. See conf/phi4-chat-template.jinja.
|
|
||||||
- --chat-template
|
|
||||||
- /config/phi4-chat-template.jinja
|
|
||||||
deploy:
|
deploy:
|
||||||
resources:
|
resources:
|
||||||
reservations:
|
reservations:
|
||||||
|
|||||||
@@ -1,24 +0,0 @@
|
|||||||
{#-
|
|
||||||
Phi-4-mini chat template — OLLAMA-MATCHING variant.
|
|
||||||
|
|
||||||
Why this exists: the official HF tokenizer template emits <|end|> after the
|
|
||||||
SYSTEM turn; Ollama's phi4 template does NOT (its system-block <|end|> is only
|
|
||||||
in the tools branch). That single boundary token regressed brokkr's R15 P02
|
|
||||||
admission eval on vLLM vs the Ollama-measured canonical (type macro-F1 -33pp)
|
|
||||||
while valid_format held at 1.0. Operator chose (2026-06-04) to make vLLM match
|
|
||||||
Ollama's leaner scaffold globally so baseline == production.
|
|
||||||
|
|
||||||
Renders (system + user, add_generation_prompt):
|
|
||||||
<|system|>{sys}<|user|>{usr}<|end|><|assistant|>
|
|
||||||
i.e. NO <|end|> after the system turn (the only delta from official).
|
|
||||||
-#}
|
|
||||||
{%- for message in messages -%}
|
|
||||||
{%- if message['role'] == 'system' -%}
|
|
||||||
{{- '<|system|>' + message['content'] -}}
|
|
||||||
{%- else -%}
|
|
||||||
{{- '<|' + message['role'] + '|>' + message['content'] + '<|end|>' -}}
|
|
||||||
{%- endif -%}
|
|
||||||
{%- endfor -%}
|
|
||||||
{%- if add_generation_prompt -%}
|
|
||||||
{{- '<|assistant|>' -}}
|
|
||||||
{%- endif -%}
|
|
||||||
Reference in New Issue
Block a user