Files
esh-pfi-infrastructure/stacks/voices-seat
Vuong Hoang 6692701571 docs(voices-seat): every voice prefixes an empty think block unless the caller disables it
Found while smoke-testing the lv-mccarthy ship. Measured live:

  default                 -> '<think>\n\n</think>\n\nThere were no horses in the road...'
  enable_thinking=false   -> 'The sun was hot on the dry riverbed and the stones were red...'

This is the Qwen3-4B-Instruct CHAT TEMPLATE, not an adapter property, so it applies
to lv-yarros, lv-bronte and lv-hemingway equally and has done since this seat went
up on 2026-09-16.

No gate number is affected: gen_beats_chat_yarros.py sets enable_thinking when the
template supports it, so every arm in every r49 gate was generated without the tags.
But a caller that does not pass chat_template_kwargs gets 17 junk characters at the
head of every passage -- and any word-count or in-band check run over that string is
counting the tags as prose. Skaldsong should be checked.
2026-09-21 18:02:56 -07:00
..

voices-seat — author voice adapters on one carrier

fv-ml1 GPU 0, port 8027. One Qwen3-4B-Instruct base; each author is a LoRA adapter named lv-<author> (lv = lang-voice). Switching voices is a request field, not a deployment.

curl http://10.251.50.54:8027/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"lv-yarros","messages":[{"role":"user","content":"Beat: ..."}]}'
#         ^^^^^^^^^ the only thing that changes between voices

voices-base serves the unadapted carrier from the same process, which is what makes an adapter-on / adapter-off comparison harness-matched.

Adding a voice

Two paths, and they are not interchangeable.

Try one now — no restart, ~0.24 s:

curl -X POST http://10.251.50.54:8027/v1/load_lora_adapter -H 'Content-Type: application/json' \
  -d '{"lora_name":"lv-hemingway","lora_path":"/adapters/lv-hemingway-4b-v1"}'
curl -X POST http://10.251.50.54:8027/v1/unload_lora_adapter -H 'Content-Type: application/json' \
  -d '{"lora_name":"lv-hemingway"}'

⚠ A runtime-loaded adapter is GONE on the next compose up -d. To make it survive, add it to --lora-modules in compose.yaml — which costs a container recreate and a full model reload (~3 min). Runtime load is for trying a voice; the compose list is what persists.

Adapters live on the host at /tank/aimodels/voice-adapters/<name>/, mounted read-only at /adapters. A rename on the host is visible inside immediately.

Reaching it through LiteLLM

Each voice is one alias entry pointing at this seat with its own model value — no new deployment, no new container, no VRAM:

- model_name: lv-yarros
  litellm_params:
    model: openai/lv-yarros
    api_base: http://10.251.50.54:8027/v1

⚠ Retiring a voice orphans scoped keys. Any LiteLLM key whose allowlist names a removed alias starts returning silent per-endpoint 403s. Audit /key/list + /key/info on every repoint — that failure has cost this fleet two and a half months before.

What it costs, measured

decode, base 143.0 tok/s (median, n=30)
decode, LoRA 108.2 tok/s (median, n=30)
LoRA overhead −24.3%, against an A-vs-A noise floor of 0.1%
resident VRAM 10,740 MiB
KV cache 10,912 tokens → 1.33x concurrency at 8k context

Arms were interleaved rather than blocked, because the co-tenants on that card take traffic this seat does not control and a block design would alias their load onto one arm.

The 24.3% is accepted deliberately. Three authors cost 8.4 GB as adapters and ~23 GB merged; six cost 9.2 GB versus ~46 GB. On a card with 1.8 GB free after this seat, that is the whole argument. If a voice ever lands on a latency path, merge that one and serve it separately.

Placement warnings

⚠ GPU 0 is now at 96.0 of 97.9 GB. This seat's 10.7 GB went into the last real gap on fv-ml1: GPU 1 has ~5.7 GB, GPU 2 has ~2.4 GB, and GPU 3 is a held reserve for a future full-card seat (flash-next alone needs 93 of 96 GiB). There is no room for another seat here without a placement decision.

⚠ --gpu-memory-utilization is a request against TOTAL VRAM that the card must already be able to honour — not a share of what is free. The first bring-up refused at 0.12 with 11.16 GiB free, and refusing was correct: it protected cyberprev, gen-small and the Parakeet STT seat rather than squeezing them.

⚠ Pin --kv-cache-memory in bytes. With it, the requested fraction predicted residency to within 40 MiB. Without it, this box has been wrong by 8–10 GB in both directions.

Support was checked, not assumed

vllm/model_executor/models/qwen3.py declares Qwen3ForCausalLM with SupportsLoRA plus packed_modules_mapping and embedding_modules. The training playbook records a LoRA refusal on a Qwen3 MoE architecture, and its lesson is that support is per-architecture, not per-family. Do not transplant this compose onto an MoE carrier without re-running that check.