# waterland-studio — watercolour render service (irv-ml1) FastAPI + vanilla-JS SPA fronting the `waterland` CLI: upload an image, get a watercolour plate and a painted-in reveal animation. Every job shells out to the CLI, which runs a fluid simulation on the **A6000**. - **Host:** irv-ml1 (10.100.79.3, WireGuard-only) · **Port:** 8410 - **URL:** http://10.100.79.3:8410/ · **Health:** `GET /api/health` - **Source:** `gitea.phasefinal.com/vh/waterland`, tracking **`main`** Handed over by `waterland-dev` on 2026-08-19, replacing a bare `nohup` that would not have survived a reboot. ## Layout — the build context is deliberately outside this directory | path | what | |---|---| | `/opt/waterland-studio/src` | checkout of `vh/waterland` @ `main` — the build context | | `/opt/docker/compose/waterland-studio/` | `compose.yaml`, `Dockerfile`, `update.sh`, `.env` | | volume `waterland-studio_waterland_studio_data` | job store (`/data`) | | volume `waterland-studio_waterland_studio_kernels` | cupy JIT cache (`/root/.cupy`) | **The checkout must NOT live under the compose directory.** `deploy-stack.sh` rsyncs `stacks//` with `--delete`, so a checkout kept beside `compose.yaml` would be destroyed by the next deploy of this stack. The Dockerfile is passed out-of-context to keep both trees clean. Refresh source + rebuild: ```bash ssh infra-ops@10.100.79.3 /opt/docker/compose/waterland-studio/update.sh ``` ## Three landmines, all of them measured rather than guessed **1. Both uv extras are load-bearing, at build AND at run.** `gpu` carries `cupy-cuda12x`; a bare `uv sync` prunes it and the renderer silently drops to the numpy path at ~21x the wall time — it does not error, it just gets slow. Worse, `studio/jobs.py` shells the renderer out as a literal `uv run waterland` with no `--extra` flags, so uv would re-sync at *runtime* and prune cupy right back out. `UV_NO_SYNC=1` stops that; `UV_OFFLINE=1` means that if the pin ever stops working the job fails loudly instead of quietly rebuilding a slower environment. **2. cupy needs CUDA *headers*, which the host never had to declare.** cupy compiles kernels at runtime through NVRTC, which needs toolkit headers — not just the driver and the runtime libs bundled in the wheel. On irv-ml1 a CUDA toolkit is installed system-wide, so the bare `nohup` process found them by accident; a slim image has none. Every render died 1.7s in with ``` RuntimeError: Failed to find CUDA headers. ``` printed *through argparse's usage banner*, which makes it read like a CLI argument bug rather than a missing toolkit — that misdirection is the reason this is written down. Fixed with `uv pip install "cupy-cuda12x[ctk]"`, which pulls the headers as wheels: a few hundred MB against ~6 GB for a `-devel` base image. It runs *after* `uv sync`, because sync prunes what it does not know about. **3. The GPU index inside the container is not the host's.** The app pins `CUDA_DEVICE_ORDER=PCI_BUS_ID` and selects `CUDA_VISIBLE_DEVICES_TARGET` (default `1`, correct on the host). Compose exposes exactly one GPU (`device_ids: ["1"]`, the A6000 in Docker's ordering), so **inside** the container that card is index **0** — hence `CUDA_VISIBLE_DEVICES_TARGET=0`. Copying the host's value selects a device that does not exist. Device 0 on the host is the 3090, which hosts the TTS zoo and must not be touched. ## Performance, measured on this host | job | wall | |---|---| | 256², animation, **cold container** | 23.3 s | | 256², animation, warm | 6.1 s | | 256², plate only (`--codec none`) | 3.9 s | | 512², plate only | 6.4 s | Warm numbers beat the 7.4 s at 256² recorded against the bare-metal process, so containerising cost nothing. The cold-vs-warm gap is **cupy's NVRTC compile**, which is why `/root/.cupy` is a volume: verified by recreating the container (fresh cache → 23.2 s first render) and then restarting it (populated cache → 6.0 s). Without that volume every restart makes the next user wait 4x and the service merely looks slow. ## Operational constraints — from waterland-dev, not inferred - **Serial by design. One replica, one card.** A render is 20–45 s of near-full GPU and the app runs a single worker thread. Two of these on the same A6000 would OOM or thrash. Throughput is a conversation about hardware, not replica count. - **No authentication, and it accepts arbitrary file uploads.** It must stay inside the LAN / WireGuard boundary. Do **not** paper over this with a proxy password — waterland-dev has offered to add a real auth layer if it ever needs wider reach. Ask. - Job store is scratch output, not source-of-truth: ~12 MB per animated job, self-evicting at 40 retained jobs (`RETAIN` in `studio/jobs.py`), so steady state is bounded around 500 MB. As of `b72425b` the index is rehydrated from disk at startup, so that bound holds **across restarts** and not merely within one process — see the resolved-finding section below for why that distinction mattered. The 4 jobs from the bare-metal instance were copied in at cutover. - Internal render timeout is 480 s, which is why the healthcheck interval is loose — an aggressive probe would measure queue depth rather than liveness. ## Source credential — `claude-bot`, read-only The repo is not anonymously readable (an unauthenticated clone 403s), so the host needs a credential to fetch. Operator granted **`claude-bot` read on `vh/waterland`** on 2026-08-19; verified scoped correctly: ``` admin: False | push: False | pull: True ``` The token lives on irv-ml1 at `/root/.config/waterland-studio/git-credentials`, mode `0600`, root-owned, and is wired as a **repo-scoped** credential helper — not a global one: ```bash git -C /opt/waterland-studio/src config credential.helper \ 'store --file=/root/.config/waterland-studio/git-credentials' ``` `.git/config` itself holds no token (verified), so the remote URL stays clean in any diff, log or backup of the checkout. The operator's `vh` site-admin token was used **only** for the initial clone and the grant itself, passed inline and never written to disk on this host. Do not reintroduce it: a site-admin credential on a GPU box is a blast radius nobody needs for a read-only fetch. ## ✅ Resolved upstream: the on-disk job store used to grow without bound **Fixed in `b72425b` and deployed here.** Kept on the record because the symptom is easy to misread and the fix has a property worth knowing about. The job index used to be **memory-only** — `JobStore._jobs` was a plain dict and nothing scanned `WATERLAND_STUDIO_DATA` at startup. Two consequences: 1. After any restart the UI and `/api/jobs` listed only jobs created *since* that restart, even though every earlier job's files were still on the volume. Cosmetic, not data loss — and it is how this was spotted: the API reported 1 job while the volume held all 16 directories (60.6 MB). 2. The real one: `RETAIN = 40` self-eviction only ever walked the in-memory dict, so directories orphaned by a restart were **never** reclaimed. The handover's "bounded around 500 MB" held within a single process lifetime; across restarts the store grew monotonically at ~12 MB per animated job. Upstream now **rehydrates the index at startup** by scanning the data dir, and deliberately *adopts* directories it cannot fully parse rather than skipping them — a directory that never enters the index is precisely the one eviction can never reclaim. Sidecar present → restored; no sidecar → adopted with dimensions read from the PNG IHDR; corrupt sidecar → inferred, no startup crash. **Neither source nor sidecar → skipped on purpose**, because adopting those would make eviction a delete-arbitrary-directories primitive pointed at this volume. Jobs left `running`/`queued` by a mid-render restart are marked `failed`, so they stop being unreclaimable phantoms that inflate `queue_depth`. Verified here after the update: `/api/jobs` went **1 → 16** against 16 directories on disk, with nothing reclaimed — correct, since 16 is under `RETAIN=40`; adoption only made them visible. The bound now holds across restarts, so no hand-pruning is needed. **Don't chase the one `failed` job in that list.** Adoption marks a job `failed` on exactly two conditions: a sidecar saying `running`/`queued`, or a directory with no `plate.png`. The renders that died 1.7 s in on the missing CUDA headers (before that fix) wrote a source and never a plate, so they land in the second branch. It is adopted history being labelled honestly, not a current fault.