The ~6.5-9 s first-call-per-bucket autotune lived in the container's writable layer and died on every recreate. 0.1.3 creates /tmp/triton-cache in the image owned by 10001 so the named volume intern-decision_triton-cache inherits a writable mount point, and compose mounts it. Bucket model PROVEN, not inferred: 2,048-token buckets, 16 up to 32,768. After one warmed call per bucket, 12 random sizes across 8k-32k were all warm (worst 2.09 s); cold entries cost 6.5-9 s. Full cold warm-up 109 s; warm re-run 17 s. scripts/intern-decision-warmup: one noul call per bucket, MAX_TOKENS from /health, two-point live calibration of the tokenizer's linear token model (a single probe overcorrects and the aim oscillates around the bucket edge), per-bucket wall times, non-zero exit on a missed bucket. Run it after an IMAGE CHANGE only; the volume carries ordinary recreates (measured: force-recreate, then a warmed 32k call answered in 2.11 s). Acceptance on 0.1.3: JevBench 202/231, hard 83/111, 0 diffs / 924; warm 32k GPU 1 peak 15,218 MiB (budget 15,220; a COLD autotune touched 15,224 once, README caveat); /decide answers. Artifacts in the acceptance dir.
55 lines
2.8 KiB
Docker
55 lines
2.8 KiB
Docker
# syntax=docker/dockerfile:1
|
|
# intern-decision-serve: Intern-Decision-4B, scored by the checkpoint's own inference.py, behind
|
|
# semif-serve's HTTP surface. Contract: intern-decision-serve.contract.md.
|
|
# docker build -t intern-decision-serve:<version> .
|
|
# Weights AND inference.py are NOT in the image: both are read from the mounted HF cache at the
|
|
# pinned revision, offline (INV-5); inference.py's sha256 is checked before it is imported (INV-3).
|
|
|
|
# The dependency manifest with the service's own version blanked to 0.0.0, so a version bump
|
|
# leaves these two files byte-identical and the ~4 GB torch/CUDA layer below stays cached.
|
|
FROM python:3.12-slim-bookworm AS deps
|
|
WORKDIR /deps
|
|
COPY pyproject.toml uv.lock ./
|
|
RUN python - <<'PY'
|
|
import re, pathlib
|
|
p = pathlib.Path("pyproject.toml")
|
|
p.write_text(re.sub(r'(?m)^version = "[^"]+"', 'version = "0.0.0"', p.read_text(), count=1))
|
|
l = pathlib.Path("uv.lock")
|
|
l.write_text(re.sub(r'(name = "intern-decision-serve"\nversion = )"[^"]+"', r'\1"0.0.0"', l.read_text(), count=1))
|
|
PY
|
|
|
|
FROM python:3.12-slim-bookworm
|
|
# The model plus Qwen3.5's fast kernels (fla + causal-conv1d): the bench stack.
|
|
ARG EXTRAS="--extra model --extra fast"
|
|
COPY --from=ghcr.io/astral-sh/uv:0.6.9 /uv /bin/uv
|
|
ENV UV_COMPILE_BYTECODE=1 UV_LINK_MODE=copy UV_PYTHON_DOWNLOADS=never
|
|
# The fast extra needs a C compiler AT RUNTIME: triton builds its CUDA driver shim on first use,
|
|
# and without gcc the warm-up dies with "Failed to find C compiler" (semif-serve, 2026-09-27).
|
|
RUN apt-get update && apt-get install -y --no-install-recommends ca-certificates \
|
|
&& if echo "$EXTRAS" | grep -q -- '--extra fast'; then \
|
|
apt-get install -y --no-install-recommends gcc libc6-dev; fi \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
WORKDIR /app
|
|
COPY --from=deps /deps/pyproject.toml /deps/uv.lock ./
|
|
RUN --mount=type=cache,target=/root/.cache/uv \
|
|
uv sync --frozen --no-dev $EXTRAS --no-install-project
|
|
COPY pyproject.toml uv.lock ./
|
|
COPY src ./src
|
|
RUN uv sync --frozen --no-dev $EXTRAS --no-editable --no-cache
|
|
RUN groupadd --system --gid 10001 intern \
|
|
&& useradd --system --uid 10001 --gid 10001 --no-create-home --shell /usr/sbin/nologin intern
|
|
USER intern
|
|
# The mount point for the Triton autotune cache volume, created HERE so a fresh named volume
|
|
# inherits intern:intern (a root-owned mount point would make every autotune write fail).
|
|
RUN mkdir -p /tmp/triton-cache
|
|
ENV PATH=/app/.venv/bin:$PATH \
|
|
HF_HOME=/hf \
|
|
HF_HUB_OFFLINE=1 \
|
|
TRANSFORMERS_OFFLINE=1 \
|
|
HF_HUB_DISABLE_TELEMETRY=1 \
|
|
TRITON_CACHE_DIR=/tmp/triton-cache \
|
|
NVIDIA_DRIVER_CAPABILITIES=compute,utility
|
|
EXPOSE 8000
|
|
# One worker (INV-2): the model and the inference lock live in this one process.
|
|
CMD ["uvicorn", "intern_decision_serve.main:app_from_env", "--factory", "--host", "0.0.0.0", "--port", "8000", "--workers", "1"]
|