Files
esh-pfi-infrastructure/services/semif-serve/Dockerfile
T
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00

57 lines
2.9 KiB
Docker

# syntax=docker/dockerfile:1
# semif-serve: SemIf (pinned commit) behind a small FastAPI service. Contract: semif-serve.contract.md.
# docker build -t semif-serve:<version> .
# Weights are NOT in the image: the pinned Qwen3.5-4B revision is read from the mounted
# HF cache, offline (INV-5).
# The dependency manifest with semif-serve's own version blanked to 0.0.0. A version bump
# then leaves these two files byte-identical, and COPY --from compares CONTENT, so the ~4 GB
# torch/CUDA install below stays cached across releases (the same problem augaman hit).
# `uv sync` keeps uv's per-package index routing (torch from the cu128 index, everything
# else from PyPI). An exported requirements.txt loses that, and then fetches triton from the
# wrong index and fails its hash check.
FROM python:3.12-slim-bookworm AS deps
WORKDIR /deps
COPY pyproject.toml uv.lock ./
RUN python - <<'EOF'
import re, pathlib
p = pathlib.Path("pyproject.toml")
p.write_text(re.sub(r'(?m)^version = "[^"]+"', 'version = "0.0.0"', p.read_text(), count=1))
l = pathlib.Path("uv.lock")
l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"', l.read_text(), count=1))
EOF
FROM python:3.12-slim-bookworm
# EXTRAS picks the optional dependency sets. Default (adopted 2026-09-27): the model plus
# Qwen3.5's fast kernels (fla + causal-conv1d). "--extra model" alone gives the reference
# PyTorch paths.
ARG EXTRAS="--extra model --extra fast"
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
# git: semif-phase1 installs from a pinned GitHub commit. The fast extra also needs a C
# compiler AT RUNTIME: triton builds its CUDA driver shim on first use, and without gcc the
# warm-up dies with "Failed to find C compiler" (seen 2026-09-27; startup failed closed).
RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \
&& if echo "$EXTRAS" | grep -q -- '--extra fast'; then \
apt-get install -y --no-install-recommends gcc libc6-dev; fi \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev $EXTRAS --no-install-project
COPY pyproject.toml uv.lock ./
COPY src ./src
RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache
RUN groupadd --system --gid 10001 semif \
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif
USER semif
ENV PATH=/app/.venv/bin:$PATH \
HF_HOME=/hf \
HF_HUB_OFFLINE=1 \
HF_HUB_DISABLE_TELEMETRY=1 \
TRITON_CACHE_DIR=/tmp/triton-cache \
NVIDIA_DRIVER_CAPABILITIES=compute,utility
EXPOSE 8000
# One worker (INV-2): the model and the inference lock live in this one process.
CMD ["uvicorn", "semif_serve.main:app_from_env", "--factory", "--host", "0.0.0.0", "--port", "8000", "--workers", "1"]