feat(nh3-ml1): LFM2.5-VL-3B (llama.cpp) + VibeVoice-ASR-Streaming-1.5B (audio.cpp) utility seats

For brokkr's dataset foundry (operator-approved 2026-09-26, relayed).
- stacks/lfm-vl-seat: llama.cpp server-cuda b11176 (digest-pinned), Q5_K_M +
  mmproj Q8_0, :8030; gateway alias lfm25-vl-3b (LiteLLM restarted, 36 s).
  Positive control exact; null control shows it describes a missing image.
- stacks/vibevoice-asr-seat: audio.cpp v0.8.2-audio8-perf-hotfix (the GGUF's
  own runtime, not vibevoice.cpp) on cuda 12.8 runtime + libgomp + libsoxr,
  sha256-pinned; :8031 direct. LibriSpeech WER 3/69, RTF 0.07-0.14; ~31 s
  cold first request.
This commit is contained in:
vh
2026-09-26 00:41:01 -07:00
parent 45484a0007
commit f2792183d4
11 changed files with 235 additions and 0 deletions
+14
View File
@@ -512,6 +512,20 @@ model_list:
model_info:
mode: rerank
# --- lfm25-vl-3b → LiquidAI LFM2.5-VL-3B (Q5_K_M + mmproj Q8_0), llama.cpp server on
# nh3-ml1 :8030 (stacks/lfm-vl-seat, 2026-09-26). Image understanding for
# brokkr's dataset foundry; send images as OpenAI image_url parts. A small VLM:
# with NO image attached it confidently describes one anyway (measured), so
# callers must check the image actually went in. Batch-grade speed (~90 tok/s). ---
- model_name: lfm25-vl-3b
litellm_params:
model: hosted_vllm/lfm25-vl-3b
api_base: http://10.100.50.80:8030/v1
api_key: os.environ/VLLM_API_KEY
model_info:
mode: chat
supports_vision: true
# --- coder-fast → Qwen2.5-Coder-1.5B (BASE), FIM code-completion seat (ana-ml2
# GPU1 :8020, vLLM; deep-research pick 2026-07-27). For Zed editor inline
# edit-predictions via the LEGACY /v1/completions endpoint with Qwen FIM