Found while smoke-testing the lv-mccarthy ship. Measured live: default -> '<think>\n\n</think>\n\nThere were no horses in the road...' enable_thinking=false -> 'The sun was hot on the dry riverbed and the stones were red...' This is the Qwen3-4B-Instruct CHAT TEMPLATE, not an adapter property, so it applies to lv-yarros, lv-bronte and lv-hemingway equally and has done since this seat went up on 2026-09-16. No gate number is affected: gen_beats_chat_yarros.py sets enable_thinking when the template supports it, so every arm in every r49 gate was generated without the tags. But a caller that does not pass chat_template_kwargs gets 17 junk characters at the head of every passage -- and any word-count or in-band check run over that string is counting the tags as prose. Skaldsong should be checked.
voices-seat — author voice adapters on one carrier
fv-ml1 GPU 0, port 8027. One Qwen3-4B-Instruct base; each author is a LoRA adapter
named lv-<author> (lv = lang-voice). Switching voices is a request field, not a
deployment.
curl http://10.251.50.54:8027/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"lv-yarros","messages":[{"role":"user","content":"Beat: ..."}]}'
# ^^^^^^^^^ the only thing that changes between voices
voices-base serves the unadapted carrier from the same process, which is what makes an
adapter-on / adapter-off comparison harness-matched.
Adding a voice
Two paths, and they are not interchangeable.
Try one now — no restart, ~0.24 s:
curl -X POST http://10.251.50.54:8027/v1/load_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name":"lv-hemingway","lora_path":"/adapters/lv-hemingway-4b-v1"}'
curl -X POST http://10.251.50.54:8027/v1/unload_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name":"lv-hemingway"}'
⚠ A runtime-loaded adapter is GONE on the next compose up -d. To make it survive, add
it to --lora-modules in compose.yaml — which costs a container recreate and a full model
reload (~3 min). Runtime load is for trying a voice; the compose list is what persists.
Adapters live on the host at /tank/aimodels/voice-adapters/<name>/, mounted read-only at
/adapters. A rename on the host is visible inside immediately.
Reaching it through LiteLLM
Each voice is one alias entry pointing at this seat with its own model value — no new
deployment, no new container, no VRAM:
- model_name: lv-yarros
litellm_params:
model: openai/lv-yarros
api_base: http://10.251.50.54:8027/v1
⚠ Retiring a voice orphans scoped keys. Any LiteLLM key whose allowlist names a removed
alias starts returning silent per-endpoint 403s. Audit /key/list + /key/info on every
repoint — that failure has cost this fleet two and a half months before.
What it costs, measured
| decode, base | 143.0 tok/s (median, n=30) |
| decode, LoRA | 108.2 tok/s (median, n=30) |
| LoRA overhead | −24.3%, against an A-vs-A noise floor of 0.1% |
| resident VRAM | 10,740 MiB |
| KV cache | 10,912 tokens → 1.33x concurrency at 8k context |
Arms were interleaved rather than blocked, because the co-tenants on that card take traffic this seat does not control and a block design would alias their load onto one arm.
The 24.3% is accepted deliberately. Three authors cost 8.4 GB as adapters and ~23 GB merged; six cost 9.2 GB versus ~46 GB. On a card with 1.8 GB free after this seat, that is the whole argument. If a voice ever lands on a latency path, merge that one and serve it separately.
Placement warnings
⚠ GPU 0 is now at 96.0 of 97.9 GB. This seat's 10.7 GB went into the last real gap on
fv-ml1: GPU 1 has ~5.7 GB, GPU 2 has ~2.4 GB, and GPU 3 is a held reserve for a future
full-card seat (flash-next alone needs 93 of 96 GiB). There is no room for another seat
here without a placement decision.
⚠ --gpu-memory-utilization is a request against TOTAL VRAM that the card must already be
able to honour — not a share of what is free. The first bring-up refused at 0.12 with
11.16 GiB free, and refusing was correct: it protected cyberprev, gen-small and the
Parakeet STT seat rather than squeezing them.
⚠ Pin --kv-cache-memory in bytes. With it, the requested fraction predicted residency
to within 40 MiB. Without it, this box has been wrong by 8–10 GB in both directions.
Support was checked, not assumed
vllm/model_executor/models/qwen3.py declares Qwen3ForCausalLM with SupportsLoRA plus
packed_modules_mapping and embedding_modules. The training playbook records a LoRA
refusal on a Qwen3 MoE architecture, and its lesson is that support is per-architecture,
not per-family. Do not transplant this compose onto an MoE carrier without re-running that
check.