Files
esh-pfi-infrastructure/stacks/parakeet-nemo/Dockerfile
T
vh de6ea32f34 feat(parakeet-nemo): speech seat moves to parakeet-unified-en under NeMo (bf16 weights)
Prime-approved switch of the fleet STT seat (fv-ml1 :8300, LiteLLM ext-stt/
whisper-1, caller talk) from the sherpa-onnx int8 seat to arm B-bf16w of the
2026-09-30 A/B (docs/pfi/parakeet-seat-ab-2026-09-30.md): p50 33/36/42/71 ms
vs the old seat's 187/308/626 measured on the same card today, WER 1.965/3.026
vs the A/B floor 1.97/3.09. All three seat defects fixed: 12-min file 200s
(windowed at 360 s after a GPU 0 OOM on one whole request; the A/B's own
long-form method), no pause truncation, no long-form dropout.

GPU 0 room: gen-small --gpu-memory-utilization 0.48 -> 0.36 (0.46 and 0.40
refuse their boot check; cyberprev+voices hold the card). Its KV is byte-
pinned, so the boot log is token-identical: 670,142 tokens / 2.56x before
and after. Seat rests 2,088 MiB; GPU 0 keeps ~1.9 GB free.

Two runtime landmines documented in the README: NeMo's attention mask is
materialised T x T even under local attention (hence the window), and
httptools 0.8.0 writes a NUL into the HTTP status line that httpx — i.e.
LiteLLM — rejects, so the image ships plain uvicorn with --http h11.

Old seat stopped, not removed: docker stop parakeet-nemo && docker start
parakeet is the rollback.

License: NVIDIA Open Model License (accepted by Prime 2026-09-30); note in
stacks/parakeet-nemo/README.md.
2026-10-01 01:32:48 -07:00

36 lines
1.8 KiB
Docker

# Parakeet ASR seat: parakeet-unified-en-0.6b under NeMo torch, bf16 weights.
# CUDA 12.8 runtime base + a uv-managed venv pinned to the A/B's proven stack
# (torch 2.8 cu128, nemo_toolkit[asr]==3.0.0; the A/B found NeMo 2.7.3 lacks this encoder's
# att_chunk_context_size, so 3.0.0 is a floor, not a preference).
# Weights are NOT baked in: /tank/aimodels/huggingface is bind-mounted read-only (see compose).
FROM nvidia/cuda:12.8.1-base-ubuntu24.04
ENV DEBIAN_FRONTEND=noninteractive \
PIP_DISABLE_PIP_VERSION_CHECK=1 \
PYTHONUNBUFFERED=1 \
HF_HUB_OFFLINE=1
RUN apt-get update && apt-get install -y --no-install-recommends \
python3 python3-venv python3-pip wget libsndfile1 ca-certificates \
&& rm -rf /var/lib/apt/lists/*
RUN python3 -m venv /opt/venv \
&& /opt/venv/bin/pip install -q uv \
&& UV_LINK_MODE=copy /opt/venv/bin/uv pip install -q --python /opt/venv/bin/python \
--index-url https://download.pytorch.org/whl/cu128 \
--extra-index-url https://pypi.org/simple \
"torch==2.8.*" "torchaudio==2.8.*" "nemo_toolkit[asr]==3.0.0" \
fastapi "uvicorn==0.53.0" python-multipart soundfile \
&& /opt/venv/bin/python -c "import nemo, torch; print('nemo', nemo.__version__, 'torch', torch.__version__, 'cuda_ok', torch.cuda.is_available())"
WORKDIR /app
COPY app.py /app/app.py
# The seat's only writable need is NeMo/HF scratch; keep it off the rootfs surprises.
ENV HOME=/tmp
EXPOSE 8000
# --http h11: the [standard] extra pulls httptools, and httptools 0.8.0 writes a NUL into the
# status line ("HTTP/1.1 200\x00OK") that h11/httpx reject. uvicorn auto-picks httptools when
# importable, so it must stay UNinstalled and the flag must stay explicit. See README.
CMD ["/opt/venv/bin/uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000", "--http", "h11"]