# voices-seat — one base carrier, N author LoRA adapters, hot-swappable. fv-ml1 GPU 0, :8027. # # NAMING: adapters are `lv-` — lv for lang-voice (operator, 2026-09-16, retiring the # `baby*` prefix: it read fine for one experiment and invites confusion across a family). # The adapter NAME is the request's `model` field, so this string is the public API of a voice. # # WHY LORA AND NOT MERGE (operator decision 2026-09-16, after measurement): the R49 line # produces a FAMILY of author voices — lv-bronte, lv-yarros (shipped), lv-hemingway (in # build) — and Skaldsong switches between them per request. Merging bakes each 264 MB adapter # into a fresh 7.6 GB model: # # 3 authors LoRA 7.6 GB + 3x264 MB ~ 8.4 GB merged ~23 GB # 6 authors LoRA 7.6 GB + 6x264 MB ~ 9.2 GB merged ~46 GB # # On a box where GPU 1 has 5.7 GB free, GPU 2 has 2.4 GB and GPU 3 is a held reserve, that # is the whole argument. A new author becomes a directory drop, not a VRAM negotiation. # # ⚠ SUPPORT WAS CHECKED, NOT ASSUMED. The training-throughput playbook records a LoRA refusal # on a Qwen3 *MoE* arch ("get_expert_mapping must be implemented") and its lesson is that # feature support is per-architecture, not per-family. Verified on the fleet's own engine # before committing: vllm/model_executor/models/qwen3.py:271 declares Qwen3ForCausalLM with # SupportsLoRA, plus packed_modules_mapping and embedding_modules. Dense Qwen3 is fine; do NOT # transplant this compose onto an MoE carrier without re-running that grep. # # ⚠ KV IS PINNED IN BYTES, deliberately, following gen-small-seat's note. On THIS box # `--gpu-memory-utilization` has been measured wrong in both directions — cyberprev at 0.40 # holds 47.1 GB (8 GB over its fraction), gen-small at 0.48 holds 36.9 (10 GB under) — so a # fraction cannot be used to place a seat beside live co-tenants. GPU 0 already carries # cyberprev, gen-small and the Parakeet STT seat; an explicit KV budget makes this seat's # footprint deterministic instead of negotiated. # # ⚠ ADAPTERS ARE MOUNTED READ-ONLY and named explicitly. A request selects a voice by putting # the adapter name in the `model` field — switching voices is a field, not a deployment. # # MEASURED on this seat, 2026-09-16, rather than taken from docs: # POST /v1/load_lora_adapter {"lora_name","lora_path"} -> 200 in 0.24 s # POST /v1/unload_lora_adapter {"lora_name"} -> 200 in 0.003 s # VRAM unchanged across the swap; container stayed healthy; no restart, no reload. # ⚠ Anything added that way is GONE on `compose up -d` unless it is ALSO listed in # --lora-modules below. Runtime load is for trying a voice; this list is what survives. # # MEASURED LORA COST on this seat (n=30 per arm, interleaved, A-vs-A floor 0.1%): # base median 143.0 tok/s # lv-yarros median 108.2 tok/s -> -24.3%, far outside the floor. # Accepted deliberately: the seat is for prose, not a latency path, and 24% buys every future # author for 264 MB instead of 7.6 GB. If a voice ever lands on a hot path, merge THAT one. # # MEASURED RESIDENCY: 10,740 MiB, against a requested 0.11 x 94.97 GiB = 10,700 MiB — a 40 MiB # miss on a box where the util fraction has been wrong by 8-10 GB in BOTH directions. Pinning # --kv-cache-memory in bytes is what makes the fraction predictive; do not remove it. # ⚠⚠ CALLERS MUST SEND `chat_template_kwargs: {"enable_thinking": false}` OR EVERY VOICE # PREFIXES AN EMPTY `` BLOCK TO ITS OUTPUT. Measured live 2026-09-21 on # lv-mccarthy, and it is a property of the Qwen3-4B-Instruct CHAT TEMPLATE rather than of any # adapter -- so it applies to lv-yarros, lv-bronte and lv-hemingway too, and has done since # this seat went up on 2026-09-16: # # default -> '\n\n\n\nThere were no horses in the road...' # enable_thinking=false -> 'The sun was hot on the dry riverbed and the stones were red...' # # The r49 gate harness passes it (gen_beats_chat_yarros.py sets enable_thinking when the # template supports it), which is why no gate number is affected -- but a caller that does NOT # pass it gets 17 junk characters at the head of every passage, and a word-count or an # in-band check run over that string is measuring the tags as prose. Check Skaldsong. name: voices-seat services: vllm-voices: image: ${VOICES_IMAGE:-vllm/vllm-openai:nightly-e9d1398d9edfd90fcc1cf783805240e3effec013} container_name: ${VOICES_CONTAINER_NAME:-vllm-voices} restart: unless-stopped ipc: host ports: - "${VOICES_PORT:-8027}:8000" volumes: - /tank/aimodels/huggingface:/hfcache - ${VOICES_MODEL:-/tank/aimodels/Qwen3-4B-Instruct}:/model:ro - /tank/aimodels/voice-adapters:/adapters:ro environment: - VLLM_API_KEY=${API_KEY:-} - VLLM_ALLOW_RUNTIME_LORA_UPDATING=${VOICES_RUNTIME_LORA:-1} command: - /model - --served-model-name - ${VOICES_SERVED_NAME:-voices-base} - --host - 0.0.0.0 - --port - "8000" - --gpu-memory-utilization - "${VOICES_GPU_MEM_UTIL:-0.12}" - --kv-cache-memory - "${VOICES_KV_CACHE_MEMORY:-2147483648}" - --max-model-len - "${VOICES_MAX_MODEL_LEN:-8192}" - --max-num-seqs - "${VOICES_MAX_NUM_SEQS:-8}" - --dtype - auto - --enable-prefix-caching - --enable-lora # ⚠ This is the max LoRAs in a single BATCH, not the number registered — 4 adapters are # listed below and that sits exactly at the default. A FIFTH voice does not need a bump to # be registered, but it will queue rather than batch alongside the other four. - --max-loras - "${VOICES_MAX_LORAS:-4}" - --max-lora-rank - "${VOICES_MAX_LORA_RANK:-32}" - --lora-modules - lv-yarros=/adapters/lv-yarros-4b-v1 # lv-bronte: +0.193 delta_cb toward held-out Brontë, 8-gram overlap 0.00 # (identical to the never-saw-it control). Shipped 2026-09-17 as additive, # reversible and safety-clean. # # ⚠ IT WAS SHIPPED AS A VOICE-AXIS FAILURE AND THAT VERDICT NO LONGER STANDS. # As run, the floor was the largest within-arm seed spread across ALL THREE # arms = 0.251, contributed entirely by ckpt925 — a candidate that was NOT # shipped — on one outlier seed. The floor rule was changed to PAIRWISE on # 2026-09-17 (pre-registered for lv-hemingway before any of its numbers # existed, on the grounds that a candidate's verdict must not depend on which # other arms happened to be generated). Re-scored under that rule, on the # SAME generations, with the same delta_cb values: # # ckpt475 (this adapter) +0.193 vs a pairwise floor of 0.091 -> PASSES, 2.1x # ckpt925 (not shipped) +0.210 vs its own spread 0.251 -> still fails # # The previous session found this defect, wrote it down, and deliberately did # NOT exploit it — choosing the floor that passes your preferred answer after # seeing the numbers is the exact failure pre-registration exists to prevent. # The rule was fixed prospectively instead; this follows from it. # See /adapters/lv-bronte-4b-v1/README.md and # persistent-memory.d/2026-09-17-lv-bronte-gate.md - lv-bronte=/adapters/lv-bronte-4b-v1 # lv-hemingway: checkpoint-850 (epoch 0.959), NOT the loss minimum at step 1750. # The two are indistinguishable on voice — 0.072 apart against a 0.113 pairwise # floor — so the tiebreak fell to the axes that resolve, and 850 wins all of them: # 2.3x tighter seed spread (0.050 vs 0.113), lower memorisation (0.07 vs 0.08), # less ran-on (0.08 vs 0.12), and half an epoch less overfit. # # VOICE ✅ +0.413 delta_cb at 6.4x the pairwise floor — the strongest result in this # line, closing 73.8% of the span between the unadapted carrier and held-out # Hemingway itself. It also clears the OLD all-arms floor (0.113), so this verdict # does not depend on the 2026-09-17 rule change. # DAMAGE ✅ ran-on +0.08 and on-beat −0.14, both inside a 0.217 floor. # # ⚠ MEMORISATION IS THE AXIS TO READ BEFORE QUOTING THIS ONE AS CLEAN. 0.07 hit-rate # against a base control of 0.00 — but that control is weak here, because base writes # summary and cannot collide with a register it does not imitate. The honest # reference is the author himself: HELD-OUT HEMINGWAY scored against the train split # collides at 0.01. So the adapter reproduces train n-grams ~7x more often than # Hemingway reproduces himself. Every one of the 19 matched runs was READ: all are # stock dialogue ("came over and sat down at the table"), max 9 words, no proper # noun, no plot, no imagery — and 9 is shorter than the 10-word run unseen Hemingway # shares with the train split by coincidence. Elevated rate, zero protectable # content. Hemingway is in copyright; lv-yarros is the in-line precedent, also in # copyright, shipped at 0.10 against a 0.07 control. Unload is 0.003 s. # See /adapters/lv-hemingway-4b-v1/README.md, # scripts/hemingway-corpus/GATE-PREREG.md (pre-registered before any generation), # persistent-memory.d/2026-09-17-lv-hemingway-gate.md - lv-hemingway=/adapters/lv-hemingway-4b-v1 # lv-mccarthy: checkpoint-300 (epoch 0.652), NOT the loss minimum at step 900. # ⭐ The two-epoch recipe is now 0 for 3, and this time the loss curve was CONFIDENT and # wrong: ckpt900 is the minimum and ckpt300 sits +15.8x the median neighbour jitter above # it, yet ckpt300 wins every axis that resolves — 3.2x vs 1.2x voice margin, 2.7x tighter # seed spread, memorisation 1.0x the author's rate vs 1.8x, ran-on 0.13 vs 0.28. # # VOICE ✅ +0.177 delta_cb at 3.2x the pairwise floor — the BEST arm of five, on both the # primary read and the pre-registered punctuation-stripped secondary (+0.128 at 2.8x), so # ~3/4 of the gain is words rather than marks. Closed 60.8% of the reachable span. # ⚠ This corpus's voice axis is DELIBERATELY harder than Brontë's or Hemingway's: the # `mccarthy` register names its own punctuation tics and the eval drives the base control # with the same prompt, so the adapter earns no delta for "emit no quotation marks". # # NOT COPIED ✅ THE CLEANEST SAFETY RESULT IN THIS LINE. 0.12 hit-rate against HELD-OUT # McCARTHY'S OWN 0.12 — indistinguishable from the author — and its longest verbatim run # (10 words) is SHORTER than the one genuinely unseen McCarthy shares with the train split # by coincidence (12). All 31 matched runs were READ: stock grammar in the commonest words # ("looked at the wolf and he looked at the boy"), and the name-shaped hits are the RENAMED # invented names. Nothing protectable. McCarthy is in copyright with a LIVING ESTATE. # ⚠ The base-unadapted control is defective and its 0.00 is unearned — base writes summary # and cannot collide with a register it does not imitate. Against base this reads as a 12x # red flag; against the author it is 1.0x. Use the author, not base. # # ⚠⚠ THE COST, AND THE GATE'S RULE IS BLIND TO IT. Axis C is ran-on only (ratified # 2026-09-15 because the carrier maxes the other terms) and this adapter passes it with 40% # headroom — but IN-BAND IS 0.65 AGAINST BASE'S 0.89 and on-beat 0.42 against 0.71. # **35% of its generations miss the requested 90-140 word band, against base's 11%.** # Base is GOOD at length here, so the adapter measurably makes direction-following worse. # ⚠ No checkpoint choice fixes this — every adapted arm is 33-48% out-of-band and ran-on is # non-monotonic in epoch (0.38 -> 0.13 -> 0.20 -> 0.28); ckpt300 is the measured minimum. # The fix, if the cost bites, is a retrain targeting length, not a different checkpoint. # IF A CALLER DEPENDS ON A WORD BUDGET, EXPECT ONE GENERATION IN THREE TO MISS IT. # Unload is 0.003 s and one compose line. # # See /adapters/lv-mccarthy-4b-v1/README.md, # scripts/mccarthy-corpus/GATE-PREREG.md (pre-registered before any generation, plus three # dated amendments — axis C as first frozen was ARITHMETICALLY UNSATISFIABLE), # scripts/mccarthy-corpus/gate-results/ (raw output, all five arms), # persistent-memory.d/2026-09-21-lv-mccarthy-gate.md - lv-mccarthy=/adapters/lv-mccarthy-4b-v1 deploy: resources: reservations: devices: - driver: nvidia device_ids: - "${VOICES_GPU_ID:-0}" capabilities: [gpu] healthcheck: test: ["CMD-SHELL", "curl -fsS http://localhost:8000/health || exit 1"] interval: 30s timeout: 5s retries: 20 start_period: 300s labels: - homepage.group=AI - Inference - homepage.name=Voices - homepage.icon=mdi-account-voice - homepage.description=Author voice adapters (LoRA) on Qwen3-4B - homepage.href=http://10.251.50.54:8027/docs networks: [tnet] networks: tnet: name: traefik-net external: true