One update.sh run on irv-ml1 carried both open upstream PRs, per the operator's green-light on the job-store fix: - #5 (464dfc2) declares cupy-cuda12x[ctk] on the gpu extra and takes uv out of the render path (sys.executable -m waterland.cli), retiring the runtime prune trap at the source. - #6 (b72425b) rehydrates the job index from the data volume at startup, fixing the unbounded store growth reported from this side. Verified after the update rather than assumed: healthy on backend cupy; /api/jobs went 1 -> 16 against 16 directories on disk, so API and volume agree for the first time; nothing wrongly reclaimed, correct since 16 is under RETAIN=40 and adoption only makes them visible; a real 256^2 plate render completes warm, so the kernel-cache volume survived the image swap. A subsequent render took both counts to 17. The image keeps its explicit [ctk] install and UV_NO_SYNC/UV_OFFLINE pins even though both are now redundant. The header requirement is a property of this slim base, not of the upstream extra, and the cost is measured rather than assumed: uv sync satisfies it first, so the line reports "Audited 1 package" and adds 0.3s to the build. The env pins are now cheap defence-in-depth against any future path that re-enters uv. Docs corrected in place: the README's upstream-finding section is now a resolved-finding record, and the two "bounded ~500 MB" claims say which commit made that bound hold across restarts rather than only within a process. Comment-side changes pushed to the live compose dir; no restart was needed for them.
9.4 KiB
[2026-08-19] waterland studio containerised on irv-ml1 — three landmines, all measured
Handover from waterland-dev over althing (thread 01M0CDRGEZWAJCEJXXMQWXV80F):
a FastAPI + SPA GPU service fronting the waterland CLI, running as a bare
nohup (PID 1283383) that would not survive a reboot. Now
stacks/waterland-studio/, restart: unless-stopped, healthy on
irv-ml1:8410. Commits a2b5b58, 8189076.
Tracking main per operator: PR #4 merged and main HEAD was exactly the
pinned 8025366, so tracking-a-moving-ref and keeping-the-pin agreed anyway.
Now deployed at b72425b (2026-08-19). The container sat on 8025366 for
a few hours after PR #5 (464dfc2) landed — deliberately, since the image's
own guards already neutralised both landmines and the project was in
wind-down. PR #6 (the job-store rehydrate, operator-green-lit) was the rebuild
with a real reason behind it, and one update.sh run carried both. Verified
end to end after the update: healthy, backend: cupy, and a real 256² plate
render completes warm — the kernel-cache volume survived the image swap.
Build context lives OUTSIDE the compose dir — on purpose
/opt/waterland-studio/src is the checkout; the Dockerfile is passed
out-of-context from /opt/docker/compose/waterland-studio/. deploy-stack.sh
rsyncs stacks/<stack>/ with --delete, so a checkout kept beside
compose.yaml would be destroyed by the next deploy of this stack. update.sh
refreshes source → rebuild → recreate → health, and is verified end to end.
Landmine 1 — both uv extras are load-bearing at BUILD and RUN
gpu carries cupy-cuda12x; a bare uv sync prunes it and the renderer
silently drops to the numpy path at ~21x wall time — it does not error, it
just gets slow. waterland-dev warned about the build side.
The runtime side is worse and was not in the handover: studio/jobs.py
shells the renderer out as a literal uv run waterland ... with no --extra
flags (cwd=WATERLAND_STUDIO_REPO). Left alone, uv re-syncs the project
mid-job to its default extras and prunes cupy back out from under a correctly
built venv. Pinned with UV_NO_SYNC=1; UV_OFFLINE=1 alongside so that if the
pin ever stops holding the job fails loudly instead of quietly rebuilding a
slower environment.
Fixed upstream in 464dfc2: the server now spawns
sys.executable -m waterland.cli directly — no resolver in the render path at
all. The pins stay anyway. They cost nothing and are now defence-in-depth:
if any future code path re-enters uv inside the container, the job fails
loudly instead of quietly dropping to the numpy backend. uv itself must stay
in the image regardless — it performs the build-time uv sync /
uv pip install, and this is a single-stage build.
Landmine 2 — cupy needs CUDA HEADERS, which the host never had to declare
Every render died 1.7s in with:
RuntimeError: Failed to find CUDA headers.
printed through argparse's usage banner, which makes it read like a CLI argument bug rather than a missing toolkit. That misdirection is the reason this is written down.
cupy compiles kernels at runtime through NVRTC, which needs the toolkit
headers — not just the driver and the runtime libs bundled in the
cupy-cuda12x wheel. irv-ml1 has a CUDA toolkit installed system-wide, so the
bare nohup process found them by accident; a slim image has none.
Fixed with uv pip install "cupy-cuda12x[ctk]" — headers as wheels, a few
hundred MB against ~6 GB for a -devel base image. It runs after
uv sync, because sync prunes what it does not know about.
Reported upstream: it is an undeclared runtime dependency of the gpu extra,
and anyone running this without a system toolkit hits it. Declared upstream
in 464dfc2 (gpu is now cupy-cuda12x[ctk]>=13). The explicit install
stays in the Dockerfile: the header requirement is a property of this
image — a slim base with no system CUDA toolkit — so it belongs in the file
that creates the problem, not inherited from an extra two repos away. It also
survives any future restructuring of the gpu extra. Cost of keeping it is
now measured, not assumed: since uv sync satisfies it first, the line
reports Audited 1 package in 49ms and adds 0.3s to the build. A no-op
that documents a non-obvious requirement is worth 0.3s. (waterland-dev
independently agreed they would keep it too.)
Landmine 3 — the GPU index inside the container is not the host's
The app pins CUDA_DEVICE_ORDER=PCI_BUS_ID and selects
CUDA_VISIBLE_DEVICES_TARGET (default 1, correct on the host, where
nvidia-smi shows A6000 at 1). Compose exposes exactly one GPU
(device_ids: ["1"], the A6000 in Docker's ordering), so inside the
container that card is index 0 ⇒ CUDA_VISIBLE_DEVICES_TARGET=0. Copying
the host's value selects a device that does not exist. Host device 0 is the
3090, which carries the TTS zoo and must not be touched.
Cold start is ~17s of NVRTC compile → /root/.cupy is a volume
| job | wall |
|---|---|
| 256² + anim, cold container | 23.3 s |
| 256² + anim, warm | 6.1 s |
256² plate only (--codec none) |
3.9 s |
| 512² plate only | 6.4 s |
Warm beats the 7.4 s recorded against the bare-metal process, so containerising cost nothing. Verified the cache volume properly: recreate (fresh cache → 23.2 s first render) then restart (populated → 6.0 s). Without it every restart makes the next user wait 4x and the service merely looks slow.
Upstream finding — the on-disk job store grows without bound
JobStore._jobs is a plain dict and nothing scans WATERLAND_STUDIO_DATA at
startup. Consequences:
- After a restart
/api/jobslists only jobs created since — cosmetic, and how this was spotted: the API reported 1 job while the volume held all 16 directories, 60.6 MB. Not data loss. - The real one:
RETAIN = 40eviction only ever iterates the in-memory dict, so directories orphaned by a restart are never reclaimed. The handover's "bounded around 500 MB" holds within a single process lifetime; across restarts the store grows monotonically at ~12 MB per animated job.
Reported to waterland-dev with evidence; not patched from the infra side — it is their code. Prune the volume by hand if it bites first.
waterland-dev confirmed it (2026-08-19) — their "bounded ~500 MB" handover
claim holds within one process lifetime and nowhere else, which on a
restart: unless-stopped service is the wrong lifetime to have bounded. They
have surfaced a startup-rehydrate fix to the operator rather than opening a
third PR during wind-down. Operator green-lit it; PR #6 merged as b72425b
and is DEPLOYED (2026-08-19).
Startup rehydrate, as recommended — and waterland-dev deliberately went
further than the framing I sent them. I had said a directory the scan cannot
parse "just does not enter the index"; they made the opposite call, because a
directory that never enters the index is exactly the one that never gets
reclaimed. That is the sharper reading and it is the reason the fix works on
this volume at all — the 16 pre-existing dirs have no sidecar. Their
adoption ladder: sidecar → restored verbatim; no sidecar → adopted with
dimensions recovered from the PNG IHDR (24-byte read, not a decode); corrupt
sidecar → degrades to inference, no startup crash; neither source nor
sidecar → skipped on purpose, since adopting it would turn eviction into a
delete-arbitrary-directories primitive pointed at this volume. Sidecar writes
go through os.replace, and job.json is excluded from ARTIFACTS so it is
unreachable via the artifact route.
They also closed a second leak I never saw, because it needs a restart
mid-render to surface: a job left running/queued in its sidecar is
non-terminal forever, and eviction skips non-terminal jobs — so it is a
phantom that is never reclaimed and queue_depth over-reports for the life of
the process. Adoption now marks those failed.
Verified on this host after the update: /api/jobs went 1 → 16 while
the volume stayed at 16 dirs / 61 MB — disk and API agree for the first time.
Nothing was reclaimed, correctly: 16 is under RETAIN=40, so adoption only
made them visible. A subsequent real render took both to 17. From here the
store is bounded across restarts, not merely within a process.
Access
Repo is not anonymously readable (a bare clone 403s). Operator granted
claude-bot read on vh/waterland — verified admin: False, push: False, pull: True. Token on irv-ml1 at
/root/.config/waterland-studio/git-credentials, 0600 root-owned, wired as a
repo-scoped credential helper; .git/config carries no token (verified),
so the remote stays clean in any diff or backup. The operator's vh
site-admin token was used only for the initial clone and the grant itself and
was never written to disk on that host — a site-admin credential on a GPU
box is a blast radius nobody needs for a read-only fetch.
Constraints honoured as stated (not inferred)
- Serial by design — one replica, one card. A render is 20–45s of near-full GPU with a single worker thread. Two on the same A6000 would OOM or thrash. Throughput is a hardware conversation, not a replica-count one.
- No authentication, arbitrary file uploads ⇒ stays inside the LAN/WireGuard boundary. Do not paper over it with a proxy password; waterland-dev offered to add a real auth layer if wider reach is ever needed.