lv-bronte is live on vllm-voices (fv-ml1 GPU0 :8027) alongside voices-base and lv-yarros. The seat lists all three; container healthy; GPU0 96092 -> 96090 MiB, so the adapter cost nothing measurable. IT DID NOT PASS ITS VOICE GATE, and the artifact says so in three places — this commit, a comment in the compose file, and a README beside the adapter on NFS — because an adapter found without its provenance will otherwise be read as a pass. VOICE FAIL +0.193 delta_cb vs base, against a 0.251 measured noise floor NOT COPIED PASS 8-gram hit-rate 0.00, longest 0 - identical to the control NO DAMAGE PASS ran-on +0.15 against a 0.400 floor Shipped on three grounds, none of them that the number was nearly good enough: it is additive (a named LoRA nobody reaches without asking for it), reversible (one compose line; hot-unload measures 0.003 s), and clean on the axis that carries actual risk - verbatim regurgitation of the source, on a public-domain corpus, measured against a positive control that saturates at 160. The voice result is UNDERPOWERED rather than absent: it closed 48% of the span from base to the same-author target and beat the control on every individual seed. The cause is structural - 81 val pairs against Hemingway's 200, from a 678k-word corpus against 994k - and neither more beats nor more seeds fixes it, because the floor is a range statistic and ranges widen with n. ckpt475 over ckpt925: indistinguishable on voice (0.017 apart), but ckpt925 has a verbatim 8-gram hit where this has none, and is 2.7x less stable seed-to-seed (0.251 vs 0.092) with a degeneracy probe showing no collapse to explain it.
voices-seat — author voice adapters on one carrier
fv-ml1 GPU 0, port 8027. One Qwen3-4B-Instruct base; each author is a LoRA adapter
named lv-<author> (lv = lang-voice). Switching voices is a request field, not a
deployment.
curl http://10.251.50.54:8027/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"lv-yarros","messages":[{"role":"user","content":"Beat: ..."}]}'
# ^^^^^^^^^ the only thing that changes between voices
voices-base serves the unadapted carrier from the same process, which is what makes an
adapter-on / adapter-off comparison harness-matched.
Adding a voice
Two paths, and they are not interchangeable.
Try one now — no restart, ~0.24 s:
curl -X POST http://10.251.50.54:8027/v1/load_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name":"lv-hemingway","lora_path":"/adapters/lv-hemingway-4b-v1"}'
curl -X POST http://10.251.50.54:8027/v1/unload_lora_adapter -H 'Content-Type: application/json' \
-d '{"lora_name":"lv-hemingway"}'
⚠ A runtime-loaded adapter is GONE on the next compose up -d. To make it survive, add
it to --lora-modules in compose.yaml — which costs a container recreate and a full model
reload (~3 min). Runtime load is for trying a voice; the compose list is what persists.
Adapters live on the host at /tank/aimodels/voice-adapters/<name>/, mounted read-only at
/adapters. A rename on the host is visible inside immediately.
Reaching it through LiteLLM
Each voice is one alias entry pointing at this seat with its own model value — no new
deployment, no new container, no VRAM:
- model_name: lv-yarros
litellm_params:
model: openai/lv-yarros
api_base: http://10.251.50.54:8027/v1
⚠ Retiring a voice orphans scoped keys. Any LiteLLM key whose allowlist names a removed
alias starts returning silent per-endpoint 403s. Audit /key/list + /key/info on every
repoint — that failure has cost this fleet two and a half months before.
What it costs, measured
| decode, base | 143.0 tok/s (median, n=30) |
| decode, LoRA | 108.2 tok/s (median, n=30) |
| LoRA overhead | −24.3%, against an A-vs-A noise floor of 0.1% |
| resident VRAM | 10,740 MiB |
| KV cache | 10,912 tokens → 1.33x concurrency at 8k context |
Arms were interleaved rather than blocked, because the co-tenants on that card take traffic this seat does not control and a block design would alias their load onto one arm.
The 24.3% is accepted deliberately. Three authors cost 8.4 GB as adapters and ~23 GB merged; six cost 9.2 GB versus ~46 GB. On a card with 1.8 GB free after this seat, that is the whole argument. If a voice ever lands on a latency path, merge that one and serve it separately.
Placement warnings
⚠ GPU 0 is now at 96.0 of 97.9 GB. This seat's 10.7 GB went into the last real gap on
fv-ml1: GPU 1 has ~5.7 GB, GPU 2 has ~2.4 GB, and GPU 3 is a held reserve for a future
full-card seat (flash-next alone needs 93 of 96 GiB). There is no room for another seat
here without a placement decision.
⚠ --gpu-memory-utilization is a request against TOTAL VRAM that the card must already be
able to honour — not a share of what is free. The first bring-up refused at 0.12 with
11.16 GiB free, and refusing was correct: it protected cyberprev, gen-small and the
Parakeet STT seat rather than squeezing them.
⚠ Pin --kv-cache-memory in bytes. With it, the requested fraction predicted residency
to within 40 MiB. Without it, this box has been wrong by 8–10 GB in both directions.
Support was checked, not assumed
vllm/model_executor/models/qwen3.py declares Qwen3ForCausalLM with SupportsLoRA plus
packed_modules_mapping and embedding_modules. The training playbook records a LoRA
refusal on a Qwen3 MoE architecture, and its lesson is that support is per-architecture,
not per-family. Do not transplant this compose onto an MoE carrier without re-running that
check.