Files
esh-pfi-infrastructure/stacks/vibevoice-asr-seat
vh f2792183d4 feat(nh3-ml1): LFM2.5-VL-3B (llama.cpp) + VibeVoice-ASR-Streaming-1.5B (audio.cpp) utility seats
For brokkr's dataset foundry (operator-approved 2026-09-26, relayed).
- stacks/lfm-vl-seat: llama.cpp server-cuda b11176 (digest-pinned), Q5_K_M +
  mmproj Q8_0, :8030; gateway alias lfm25-vl-3b (LiteLLM restarted, 36 s).
  Positive control exact; null control shows it describes a missing image.
- stacks/vibevoice-asr-seat: audio.cpp v0.8.2-audio8-perf-hotfix (the GGUF's
  own runtime, not vibevoice.cpp) on cuda 12.8 runtime + libgomp + libsoxr,
  sha256-pinned; :8031 direct. LibriSpeech WER 3/69, RTF 0.07-0.14; ~31 s
  cold first request.
2026-09-26 00:41:01 -07:00
..

vibevoice-asr-seat

Microsoft VibeVoice-ASR-Streaming-1.5B (Q4_K) on nh3-ml1, served by audio.cpp (0xShug0/audio.cpp v0.8.2-audio8-perf-hotfix, Linux CUDA 12.8 build) on :8031, with direct access only. A utility seat for brokkr's dataset foundry (speech → text with inline speaker labels). Operator-approved 2026-09-26, relayed by brokkr-smithy-dev.

  • POST /v1/audio/transcriptions: multipart file=@x.wav, model=vibevoice-asr-streaming-1.5b. Returns {"text": " \n Speaker 0:...", "timing": {"rtf": ...}}.
  • POST /v1/audio/transcriptions/live?model=…&sample_rate=16000&channels=1&sample_format=s16le takes chunked raw PCM.
  • GET /v1/models, GET /health.

Why audio.cpp and not vibevoice.cpp or llama.cpp. The christopherthompson81/VibeVoice-ASR-Streaming-1.5B-GGUF files are audio.cpp packages with their sidecars embedded, per the model card, so no separate tokenizer is needed. The release binary expects the CUDA 12 runtime libraries, so Dockerfile puts it on nvidia/cuda:12.8.1-runtime with a SHA-256-pinned download, plus libgomp1 and libsoxr0. Without libsoxr it falls back to linear resampling (16k→24k), and on one clip that turned "cutter" into "country".

Checks, 2026-09-26. On the 4 LibriSpeech clips shipped with audio.cpp, WER is 3/69 = 4.35%, identical across 3 reps. Two of the three errors are "I'm" vs "I am" normalization. RTF is 0.07–0.14. ⚠ The first request after a start takes ~31 s (CUDA graph warmup); later ones take 0.3–1 s. VRAM ~2.0 GB.

The model card recommends Q8_0 (3.3 GB; its CUDA WER is 5.80% vs Q4_K's 7.25%, on 69 words). To switch, change path in conf/server.json, download the file, and recreate the container.