lv-bronte is live on vllm-voices (fv-ml1 GPU0 :8027) alongside voices-base and lv-yarros. The seat lists all three; container healthy; GPU0 96092 -> 96090 MiB, so the adapter cost nothing measurable. IT DID NOT PASS ITS VOICE GATE, and the artifact says so in three places — this commit, a comment in the compose file, and a README beside the adapter on NFS — because an adapter found without its provenance will otherwise be read as a pass. VOICE FAIL +0.193 delta_cb vs base, against a 0.251 measured noise floor NOT COPIED PASS 8-gram hit-rate 0.00, longest 0 - identical to the control NO DAMAGE PASS ran-on +0.15 against a 0.400 floor Shipped on three grounds, none of them that the number was nearly good enough: it is additive (a named LoRA nobody reaches without asking for it), reversible (one compose line; hot-unload measures 0.003 s), and clean on the axis that carries actual risk - verbatim regurgitation of the source, on a public-domain corpus, measured against a positive control that saturates at 160. The voice result is UNDERPOWERED rather than absent: it closed 48% of the span from base to the same-author target and beat the control on every individual seed. The cause is structural - 81 val pairs against Hemingway's 200, from a 678k-word corpus against 994k - and neither more beats nor more seeds fixes it, because the floor is a range statistic and ranges widen with n. ckpt475 over ckpt925: indistinguishable on voice (0.017 apart), but ckpt925 has a verbatim 8-gram hit where this has none, and is 2.7x less stable seed-to-seed (0.251 vs 0.092) with a degeneracy probe showing no collapse to explain it.
129 lines
5.9 KiB
YAML
129 lines
5.9 KiB
YAML
# voices-seat — one base carrier, N author LoRA adapters, hot-swappable. fv-ml1 GPU 0, :8027.
|
|
#
|
|
# NAMING: adapters are `lv-<author>` — lv for lang-voice (operator, 2026-09-16, retiring the
|
|
# `baby*` prefix: it read fine for one experiment and invites confusion across a family).
|
|
# The adapter NAME is the request's `model` field, so this string is the public API of a voice.
|
|
#
|
|
# WHY LORA AND NOT MERGE (operator decision 2026-09-16, after measurement): the R49 line
|
|
# produces a FAMILY of author voices — lv-bronte, lv-yarros (shipped), lv-hemingway (in
|
|
# build) — and Skaldsong switches between them per request. Merging bakes each 264 MB adapter
|
|
# into a fresh 7.6 GB model:
|
|
#
|
|
# 3 authors LoRA 7.6 GB + 3x264 MB ~ 8.4 GB merged ~23 GB
|
|
# 6 authors LoRA 7.6 GB + 6x264 MB ~ 9.2 GB merged ~46 GB
|
|
#
|
|
# On a box where GPU 1 has 5.7 GB free, GPU 2 has 2.4 GB and GPU 3 is a held reserve, that
|
|
# is the whole argument. A new author becomes a directory drop, not a VRAM negotiation.
|
|
#
|
|
# ⚠ SUPPORT WAS CHECKED, NOT ASSUMED. The training-throughput playbook records a LoRA refusal
|
|
# on a Qwen3 *MoE* arch ("get_expert_mapping must be implemented") and its lesson is that
|
|
# feature support is per-architecture, not per-family. Verified on the fleet's own engine
|
|
# before committing: vllm/model_executor/models/qwen3.py:271 declares Qwen3ForCausalLM with
|
|
# SupportsLoRA, plus packed_modules_mapping and embedding_modules. Dense Qwen3 is fine; do NOT
|
|
# transplant this compose onto an MoE carrier without re-running that grep.
|
|
#
|
|
# ⚠ KV IS PINNED IN BYTES, deliberately, following gen-small-seat's note. On THIS box
|
|
# `--gpu-memory-utilization` has been measured wrong in both directions — cyberprev at 0.40
|
|
# holds 47.1 GB (8 GB over its fraction), gen-small at 0.48 holds 36.9 (10 GB under) — so a
|
|
# fraction cannot be used to place a seat beside live co-tenants. GPU 0 already carries
|
|
# cyberprev, gen-small and the Parakeet STT seat; an explicit KV budget makes this seat's
|
|
# footprint deterministic instead of negotiated.
|
|
#
|
|
# ⚠ ADAPTERS ARE MOUNTED READ-ONLY and named explicitly. A request selects a voice by putting
|
|
# the adapter name in the `model` field — switching voices is a field, not a deployment.
|
|
#
|
|
# MEASURED on this seat, 2026-09-16, rather than taken from docs:
|
|
# POST /v1/load_lora_adapter {"lora_name","lora_path"} -> 200 in 0.24 s
|
|
# POST /v1/unload_lora_adapter {"lora_name"} -> 200 in 0.003 s
|
|
# VRAM unchanged across the swap; container stayed healthy; no restart, no reload.
|
|
# ⚠ Anything added that way is GONE on `compose up -d` unless it is ALSO listed in
|
|
# --lora-modules below. Runtime load is for trying a voice; this list is what survives.
|
|
#
|
|
# MEASURED LORA COST on this seat (n=30 per arm, interleaved, A-vs-A floor 0.1%):
|
|
# base median 143.0 tok/s
|
|
# lv-yarros median 108.2 tok/s -> -24.3%, far outside the floor.
|
|
# Accepted deliberately: the seat is for prose, not a latency path, and 24% buys every future
|
|
# author for 264 MB instead of 7.6 GB. If a voice ever lands on a hot path, merge THAT one.
|
|
#
|
|
# MEASURED RESIDENCY: 10,740 MiB, against a requested 0.11 x 94.97 GiB = 10,700 MiB — a 40 MiB
|
|
# miss on a box where the util fraction has been wrong by 8-10 GB in BOTH directions. Pinning
|
|
# --kv-cache-memory in bytes is what makes the fraction predictive; do not remove it.
|
|
name: voices-seat
|
|
|
|
services:
|
|
vllm-voices:
|
|
image: ${VOICES_IMAGE:-vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013}
|
|
container_name: ${VOICES_CONTAINER_NAME:-vllm-voices}
|
|
restart: unless-stopped
|
|
ipc: host
|
|
ports:
|
|
- "${VOICES_PORT:-8027}:8000"
|
|
volumes:
|
|
- /tank/aimodels/huggingface:/hfcache
|
|
- ${VOICES_MODEL:-/tank/aimodels/Qwen3-4B-Instruct}:/model:ro
|
|
- /tank/aimodels/voice-adapters:/adapters:ro
|
|
environment:
|
|
- VLLM_API_KEY=${API_KEY:-}
|
|
- VLLM_ALLOW_RUNTIME_LORA_UPDATING=${VOICES_RUNTIME_LORA:-1}
|
|
command:
|
|
- /model
|
|
- --served-model-name
|
|
- ${VOICES_SERVED_NAME:-voices-base}
|
|
- --host
|
|
- 0.0.0.0
|
|
- --port
|
|
- "8000"
|
|
- --gpu-memory-utilization
|
|
- "${VOICES_GPU_MEM_UTIL:-0.12}"
|
|
- --kv-cache-memory
|
|
- "${VOICES_KV_CACHE_MEMORY:-2147483648}"
|
|
- --max-model-len
|
|
- "${VOICES_MAX_MODEL_LEN:-8192}"
|
|
- --max-num-seqs
|
|
- "${VOICES_MAX_NUM_SEQS:-8}"
|
|
- --dtype
|
|
- auto
|
|
- --enable-prefix-caching
|
|
- --enable-lora
|
|
- --max-loras
|
|
- "${VOICES_MAX_LORAS:-4}"
|
|
- --max-lora-rank
|
|
- "${VOICES_MAX_LORA_RANK:-32}"
|
|
- --lora-modules
|
|
- lv-yarros=/adapters/lv-yarros-4b-v1
|
|
# ⚠ lv-bronte did NOT pass its voice gate: +0.193 delta_cb toward held-out
|
|
# Brontë against a 0.251 measured noise floor. Shipped anyway because it is
|
|
# additive, reversible, and clean on the safety axis (8-gram overlap 0.00,
|
|
# identical to the never-saw-it control). The effect is underpowered, not
|
|
# absent — it closed 48% of the achievable span and beat base on every seed.
|
|
# Do not cite it as evidence pair-SFT works for this author.
|
|
# See /adapters/lv-bronte-4b-v1/README.md and
|
|
# persistent-memory.d/2026-09-17-lv-bronte-gate.md
|
|
- lv-bronte=/adapters/lv-bronte-4b-v1
|
|
deploy:
|
|
resources:
|
|
reservations:
|
|
devices:
|
|
- driver: nvidia
|
|
device_ids:
|
|
- "${VOICES_GPU_ID:-0}"
|
|
capabilities: [gpu]
|
|
healthcheck:
|
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8000/health || exit 1"]
|
|
interval: 30s
|
|
timeout: 5s
|
|
retries: 20
|
|
start_period: 300s
|
|
labels:
|
|
- homepage.group=AI - Inference
|
|
- homepage.name=Voices
|
|
- homepage.icon=mdi-account-voice
|
|
- homepage.description=Author voice adapters (LoRA) on Qwen3-4B
|
|
- homepage.href=http://10.251.50.54:8027/docs
|
|
networks: [tnet]
|
|
|
|
networks:
|
|
tnet:
|
|
name: traefik-net
|
|
external: true
|