Files
esh-pfi-infrastructure/services/semif-serve/Dockerfile
T
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00

48 lines
2.3 KiB
Docker

# syntax=docker/dockerfile:1
# semif-serve: SemIf (pinned commit) behind a small FastAPI service. Contract: semif-serve.contract.md.
# docker build -t semif-serve:<version> .
# Weights are NOT in the image: the pinned Qwen3.5-4B revision is read from the mounted
# HF cache, offline (INV-5).
# The dependency manifest with semif-serve's own version blanked to 0.0.0. A version bump
# then leaves these two files byte-identical, and COPY --from compares CONTENT, so the ~4 GB
# torch/CUDA install below stays cached across releases (the same problem augaman hit).
# `uv sync` keeps uv's per-package index routing (torch from the cu128 index, everything
# else from PyPI). An exported requirements.txt loses that, and then fetches triton from the
# wrong index and fails its hash check.
FROM python:3.12-slim-bookworm AS deps
WORKDIR /deps
COPY pyproject.toml uv.lock ./
RUN python - <<'EOF'
import re, pathlib
p = pathlib.Path("pyproject.toml")
p.write_text(re.sub(r'(?m)^version = "[^"]+"', 'version = "0.0.0"', p.read_text(), count=1))
l = pathlib.Path("uv.lock")
l.write_text(re.sub(r'(name = "semif-serve"\nversion = )"[^"]+"', r'\1"0.0.0"', l.read_text(), count=1))
EOF
FROM python:3.12-slim-bookworm
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
# git: semif-phase1 installs from a pinned GitHub commit.
RUN apt-get update && apt-get install -y --no-install-recommends git ca-certificates \
&& rm -rf /var/lib/apt/lists/*
WORKDIR /app
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
RUN --mount=type=cache,target=/root/.cache/uv \
uv sync --frozen --no-dev --extra model --no-install-project
COPY pyproject.toml uv.lock ./
COPY src ./src
RUN uv sync --frozen --no-dev --extra model --no-editable --no-cache
RUN groupadd --system --gid 10001 semif \
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin semif
USER semif
ENV PATH=/app/.venv/bin:$PATH \
HF_HOME=/hf \
HF_HUB_OFFLINE=1 \
HF_HUB_DISABLE_TELEMETRY=1 \
NVIDIA_DRIVER_CAPABILITIES=compute,utility
EXPOSE 8000
# One worker (INV-2): the model and the inference lock live in this one process.
CMD ["uvicorn", "semif_serve.main:app_from_env", "--factory", "--host", "0.0.0.0", "--port", "8000", "--workers", "1"]