Files
esh-pfi-infrastructure/stacks/voices-seat
vh 2e9b118e70 lv-bronte: the voice axis passes under the corrected floor rule — amended, not rewritten
lv-bronte shipped 2026-09-17 with a FAILED voice axis written into its compose comment,
its NFS README and its gate record. That verdict no longer stands, and this records the
correction in all three places without deleting what they said.

The floor rule is now pairwise (commit 0bb4938, pre-registered for lv-hemingway before
any Hemingway number existed). Re-scoring the SAME 360 generations, no re-run, no changed
delta_cb:

  ckpt475 (shipped)      +0.193  vs pairwise floor 0.091  -> PASS, 2.1x
  ckpt925 (not shipped)  +0.210  vs its own spread 0.251  -> still fails

As run, the floor was 0.251 for every candidate, contributed entirely by ckpt925's single
outlier seed — a candidate nobody was shipping failed the one that was.

Why this is not a threshold chosen to produce a verdict: the previous session found the
defect, wrote it into this very file, and deliberately declined to act on it. The rule was
changed prospectively on an argument independent of the answer — the sampling variability
of a difference A-B depends on A and B, not on a third arm C. voice_distance.py now prints
both floors and flags disagreement so neither can be quoted without the other.

What changes for a reader: the sensitivity floor is 0.091 rather than 0.251, and "do not
cite lv-bronte as evidence pair-SFT works for this author" is withdrawn. What does not
change: NOT-COPIED and NO-DAMAGE as recorded, ckpt475 over ckpt925 for the same reasons,
and the two-epoch recipe still not transferring to Brontë.

Amendments are append-only in all three artifacts. The on-host compose is unchanged so far
— this edit is comment-only and will ride with the next real deploy rather than triggering
a model reload for a comment.
2026-09-17 02:27:52 -07:00
..

voices-seat — author voice adapters on one carrier

fv-ml1 GPU 0, port 8027. One Qwen3-4B-Instruct base; each author is a LoRA adapter named lv-<author> (lv = lang-voice). Switching voices is a request field, not a deployment.

curl http://10.251.50.54:8027/v1/chat/completions -H 'Content-Type: application/json' \
  -d '{"model":"lv-yarros","messages":[{"role":"user","content":"Beat: ..."}]}'
#         ^^^^^^^^^ the only thing that changes between voices

voices-base serves the unadapted carrier from the same process, which is what makes an adapter-on / adapter-off comparison harness-matched.

Adding a voice

Two paths, and they are not interchangeable.

Try one now — no restart, ~0.24 s:

curl -X POST http://10.251.50.54:8027/v1/load_lora_adapter -H 'Content-Type: application/json' \
  -d '{"lora_name":"lv-hemingway","lora_path":"/adapters/lv-hemingway-4b-v1"}'
curl -X POST http://10.251.50.54:8027/v1/unload_lora_adapter -H 'Content-Type: application/json' \
  -d '{"lora_name":"lv-hemingway"}'

⚠ A runtime-loaded adapter is GONE on the next compose up -d. To make it survive, add it to --lora-modules in compose.yaml — which costs a container recreate and a full model reload (~3 min). Runtime load is for trying a voice; the compose list is what persists.

Adapters live on the host at /tank/aimodels/voice-adapters/<name>/, mounted read-only at /adapters. A rename on the host is visible inside immediately.

Reaching it through LiteLLM

Each voice is one alias entry pointing at this seat with its own model value — no new deployment, no new container, no VRAM:

- model_name: lv-yarros
  litellm_params:
    model: openai/lv-yarros
    api_base: http://10.251.50.54:8027/v1

⚠ Retiring a voice orphans scoped keys. Any LiteLLM key whose allowlist names a removed alias starts returning silent per-endpoint 403s. Audit /key/list + /key/info on every repoint — that failure has cost this fleet two and a half months before.

What it costs, measured

decode, base 143.0 tok/s (median, n=30)
decode, LoRA 108.2 tok/s (median, n=30)
LoRA overhead −24.3%, against an A-vs-A noise floor of 0.1%
resident VRAM 10,740 MiB
KV cache 10,912 tokens → 1.33x concurrency at 8k context

Arms were interleaved rather than blocked, because the co-tenants on that card take traffic this seat does not control and a block design would alias their load onto one arm.

The 24.3% is accepted deliberately. Three authors cost 8.4 GB as adapters and ~23 GB merged; six cost 9.2 GB versus ~46 GB. On a card with 1.8 GB free after this seat, that is the whole argument. If a voice ever lands on a latency path, merge that one and serve it separately.

Placement warnings

⚠ GPU 0 is now at 96.0 of 97.9 GB. This seat's 10.7 GB went into the last real gap on fv-ml1: GPU 1 has ~5.7 GB, GPU 2 has ~2.4 GB, and GPU 3 is a held reserve for a future full-card seat (flash-next alone needs 93 of 96 GiB). There is no room for another seat here without a placement decision.

⚠ --gpu-memory-utilization is a request against TOTAL VRAM that the card must already be able to honour — not a share of what is free. The first bring-up refused at 0.12 with 11.16 GiB free, and refusing was correct: it protected cyberprev, gen-small and the Parakeet STT seat rather than squeezing them.

⚠ Pin --kv-cache-memory in bytes. With it, the requested fraction predicted residency to within 40 MiB. Without it, this box has been wrong by 8–10 GB in both directions.

Support was checked, not assumed

vllm/model_executor/models/qwen3.py declares Qwen3ForCausalLM with SupportsLoRA plus packed_modules_mapping and embedding_modules. The training playbook records a LoRA refusal on a Qwen3 MoE architecture, and its lesson is that support is per-architecture, not per-family. Do not transplant this compose onto an MoE carrier without re-running that check.