92 lines
4.0 KiB
Markdown
92 lines
4.0 KiB
Markdown
# voices-seat — author voice adapters on one carrier
|
||
|
||
`fv-ml1` GPU 0, port **8027**. One Qwen3-4B-Instruct base; each author is a LoRA adapter
|
||
named `lv-<author>` (**lv = lang-voice**). Switching voices is a request field, not a
|
||
deployment.
|
||
|
||
```bash
|
||
curl http://10.251.50.54:8027/v1/chat/completions -H 'Content-Type: application/json' \
|
||
-d '{"model":"lv-yarros","messages":[{"role":"user","content":"Beat: ..."}]}'
|
||
# ^^^^^^^^^ the only thing that changes between voices
|
||
```
|
||
|
||
`voices-base` serves the unadapted carrier from the same process, which is what makes an
|
||
adapter-on / adapter-off comparison harness-matched.
|
||
|
||
## Adding a voice
|
||
|
||
Two paths, and they are not interchangeable.
|
||
|
||
**Try one now — no restart, ~0.24 s:**
|
||
|
||
```bash
|
||
curl -X POST http://10.251.50.54:8027/v1/load_lora_adapter -H 'Content-Type: application/json' \
|
||
-d '{"lora_name":"lv-hemingway","lora_path":"/adapters/lv-hemingway-4b-v1"}'
|
||
curl -X POST http://10.251.50.54:8027/v1/unload_lora_adapter -H 'Content-Type: application/json' \
|
||
-d '{"lora_name":"lv-hemingway"}'
|
||
```
|
||
|
||
⚠ **A runtime-loaded adapter is GONE on the next `compose up -d`.** To make it survive, add
|
||
it to `--lora-modules` in `compose.yaml` — which costs a container recreate and a full model
|
||
reload (~3 min). Runtime load is for *trying* a voice; the compose list is what persists.
|
||
|
||
Adapters live on the host at `/tank/aimodels/voice-adapters/<name>/`, mounted read-only at
|
||
`/adapters`. A rename on the host is visible inside immediately.
|
||
|
||
## Reaching it through LiteLLM
|
||
|
||
Each voice is one alias entry pointing at this seat with its own `model` value — no new
|
||
deployment, no new container, no VRAM:
|
||
|
||
```yaml
|
||
- model_name: lv-yarros
|
||
litellm_params:
|
||
model: openai/lv-yarros
|
||
api_base: http://10.251.50.54:8027/v1
|
||
```
|
||
|
||
⚠ **Retiring a voice orphans scoped keys.** Any LiteLLM key whose allowlist names a removed
|
||
alias starts returning silent per-endpoint 403s. Audit `/key/list` + `/key/info` on every
|
||
repoint — that failure has cost this fleet two and a half months before.
|
||
|
||
## What it costs, measured
|
||
|
||
| | |
|
||
|---|---|
|
||
| decode, base | **143.0 tok/s** (median, n=30) |
|
||
| decode, LoRA | **108.2 tok/s** (median, n=30) |
|
||
| **LoRA overhead** | **−24.3%**, against an A-vs-A noise floor of **0.1%** |
|
||
| resident VRAM | **10,740 MiB** |
|
||
| KV cache | 10,912 tokens → 1.33x concurrency at 8k context |
|
||
|
||
Arms were interleaved rather than blocked, because the co-tenants on that card take traffic
|
||
this seat does not control and a block design would alias their load onto one arm.
|
||
|
||
**The 24.3% is accepted deliberately.** Three authors cost 8.4 GB as adapters and ~23 GB
|
||
merged; six cost 9.2 GB versus ~46 GB. On a card with 1.8 GB free after this seat, that is
|
||
the whole argument. If a voice ever lands on a latency path, merge that one and serve it
|
||
separately.
|
||
|
||
## Placement warnings
|
||
|
||
⚠ **GPU 0 is now at 96.0 of 97.9 GB.** This seat's 10.7 GB went into the last real gap on
|
||
fv-ml1: GPU 1 has ~5.7 GB, GPU 2 has ~2.4 GB, and GPU 3 is a held reserve for a future
|
||
full-card seat (`flash-next` alone needs 93 of 96 GiB). **There is no room for another seat
|
||
here without a placement decision.**
|
||
|
||
⚠ **`--gpu-memory-utilization` is a request against TOTAL VRAM that the card must already be
|
||
able to honour** — not a share of what is free. The first bring-up refused at 0.12 with
|
||
11.16 GiB free, and refusing was correct: it protected `cyberprev`, `gen-small` and the
|
||
Parakeet STT seat rather than squeezing them.
|
||
|
||
⚠ **Pin `--kv-cache-memory` in bytes.** With it, the requested fraction predicted residency
|
||
to within 40 MiB. Without it, this box has been wrong by 8–10 GB in *both* directions.
|
||
|
||
## Support was checked, not assumed
|
||
|
||
`vllm/model_executor/models/qwen3.py` declares `Qwen3ForCausalLM` with `SupportsLoRA` plus
|
||
`packed_modules_mapping` and `embedding_modules`. The training playbook records a LoRA
|
||
refusal on a Qwen3 **MoE** architecture, and its lesson is that support is per-architecture,
|
||
not per-family. **Do not transplant this compose onto an MoE carrier without re-running that
|
||
check.**
|