feat(nh3-ml1): LFM2.5-VL-3B (llama.cpp) + VibeVoice-ASR-Streaming-1.5B (audio.cpp) utility seats
For brokkr's dataset foundry (operator-approved 2026-09-26, relayed). - stacks/lfm-vl-seat: llama.cpp server-cuda b11176 (digest-pinned), Q5_K_M + mmproj Q8_0, :8030; gateway alias lfm25-vl-3b (LiteLLM restarted, 36 s). Positive control exact; null control shows it describes a missing image. - stacks/vibevoice-asr-seat: audio.cpp v0.8.2-audio8-perf-hotfix (the GGUF's own runtime, not vibevoice.cpp) on cuda 12.8 runtime + libgomp + libsoxr, sha256-pinned; :8031 direct. LibriSpeech WER 3/69, RTF 0.07-0.14; ~31 s cold first request.
This commit is contained in:
@@ -0,0 +1,30 @@
|
||||
# vibevoice-asr-seat
|
||||
|
||||
**Microsoft VibeVoice-ASR-Streaming-1.5B** (Q4_K) on **nh3-ml1**, served by
|
||||
**audio.cpp** (`0xShug0/audio.cpp` v0.8.2-audio8-perf-hotfix, Linux CUDA 12.8
|
||||
build) on `:8031`, with direct access only. A utility seat for brokkr's dataset
|
||||
foundry (speech → text with inline speaker labels). Operator-approved 2026-09-26,
|
||||
relayed by brokkr-smithy-dev.
|
||||
|
||||
- `POST /v1/audio/transcriptions`: multipart `file=@x.wav`, `model=vibevoice-asr-streaming-1.5b`.
|
||||
Returns `{"text": " \n Speaker 0:...", "timing": {"rtf": ...}}`.
|
||||
- `POST /v1/audio/transcriptions/live?model=…&sample_rate=16000&channels=1&sample_format=s16le`
|
||||
takes chunked raw PCM.
|
||||
- `GET /v1/models`, `GET /health`.
|
||||
|
||||
**Why audio.cpp and not vibevoice.cpp or llama.cpp.** The
|
||||
`christopherthompson81/VibeVoice-ASR-Streaming-1.5B-GGUF` files are audio.cpp
|
||||
packages with their sidecars embedded, per the model card, so no separate
|
||||
tokenizer is needed. The release binary expects the CUDA 12 runtime libraries,
|
||||
so `Dockerfile` puts it on `nvidia/cuda:12.8.1-runtime` with a SHA-256-pinned
|
||||
download, plus `libgomp1` and **`libsoxr0`**. Without libsoxr it falls back to
|
||||
linear resampling (16k→24k), and on one clip that turned "cutter" into "country".
|
||||
|
||||
**Checks, 2026-09-26.** On the 4 LibriSpeech clips shipped with audio.cpp, WER
|
||||
is 3/69 = 4.35%, identical across 3 reps. Two of the three errors are "I'm" vs
|
||||
"I am" normalization. RTF is 0.07–0.14. ⚠ **The first request after a start
|
||||
takes ~31 s** (CUDA graph warmup); later ones take 0.3–1 s. VRAM ~2.0 GB.
|
||||
|
||||
The model card recommends Q8_0 (3.3 GB; its CUDA WER is 5.80% vs Q4_K's 7.25%,
|
||||
on 69 words). To switch, change `path` in `conf/server.json`, download the file,
|
||||
and recreate the container.
|
||||
Reference in New Issue
Block a user