voxtral: switch to vllm-omni serve --omni with stage config — Voxtral is a multi-stage pipeline

Fourth attempt finally found the right invocation. Voxtral is a
two-stage TTS pipeline (language_model → acoustic_transformer →
audio output), not a flat MistralForCausalLM. Standard `vllm serve`
errored with "no module named 'acoustic_transformer'" because it
loads the model as a vanilla Mistral causal LM.

Pattern from /workspace/vllm-omni/examples/online_serving/
qwen3_tts/run_server.sh (closest in-image analog):

  vllm-omni serve <MODEL> \
    --stage-configs-path vllm_omni/model_executor/stage_configs/voxtral_tts.yaml \
    --host 0.0.0.0 --port 8000 \
    --gpu-memory-utilization 0.45 \
    --trust-remote-code --omni

Key differences from previous attempt:
  * `vllm-omni` binary, not `vllm`
  * `--omni` flag activates multi-stage pipeline
  * `--stage-configs-path` points at the bundled YAML that maps
    stages to GPU + scheduler + worker classes
  * Dropped --load-format/--tokenizer-mode/--config-format=mistral
    flags — the stage config handles tokenizer_mode internally
  * --trust-remote-code is required for the acoustic_transformer
    custom code path

Default .env.example now: GPU 0 (3090) with util 0.45 (~10.6 GB
target on 24 GB GPU). The A6000 is fully booked by Fish s2-pro.
This commit is contained in:
vh
2026-04-28 00:17:21 -07:00
parent 0304464b7d
commit fe01f73d84
2 changed files with 26 additions and 18 deletions
+16 -10
View File
@@ -41,19 +41,25 @@ services:
# container init expected --model=... as argv[0]. Set entrypoint
# to `vllm serve` (the standard CLI) and pass model as positional
# + tuning flags via command.
entrypoint: ["vllm", "serve"]
# Voxtral TTS is a STAGE-BASED pipeline (language_model →
# acoustic_transformer → audio output), not a flat
# MistralForCausalLM. vllm-omni's `--omni` mode + a stage config
# YAML drives this. The standard `vllm serve` errors with
# "no module named 'acoustic_transformer'" because it tries to
# load Voxtral as a vanilla Mistral causal LM.
#
# Pattern lifted from /workspace/vllm-omni/examples/online_serving/
# qwen3_tts/run_server.sh (closest analog example in the image).
# Stage config path is relative to WORKDIR=/workspace/vllm-omni.
entrypoint: ["vllm-omni", "serve"]
command:
- "${VOXTRAL_MODEL:-mistralai/Voxtral-4B-TTS-2603}"
- "--stage-configs-path=vllm_omni/model_executor/stage_configs/voxtral_tts.yaml"
- "--host=0.0.0.0"
- "--port=8000"
- "--dtype=bfloat16"
- "--gpu-memory-utilization=${VOXTRAL_GPU_UTIL:-0.85}"
# Voxtral uses the native Mistral model format (params.json,
# tekken.json tokenizer, consolidated.safetensors) — NOT HF
# transformers format. vLLM rejects it with "ensure presence
# of params.json" unless these three flags are set.
- "--load-format=mistral"
- "--tokenizer-mode=mistral"
- "--config-format=mistral"
- "--gpu-memory-utilization=${VOXTRAL_GPU_UTIL:-0.45}"
- "--trust-remote-code"
- "--omni"
healthcheck:
# vLLM-Omni exposes /health for liveness + /v1/models for readiness.
# /health 200 means the server's listening; /v1/models 200 means