feat(waterland-studio): deploy b72425b — all three upstream findings fixed
One update.sh run on irv-ml1 carried both open upstream PRs, per the operator's green-light on the job-store fix: - #5 (464dfc2) declares cupy-cuda12x[ctk] on the gpu extra and takes uv out of the render path (sys.executable -m waterland.cli), retiring the runtime prune trap at the source. - #6 (b72425b) rehydrates the job index from the data volume at startup, fixing the unbounded store growth reported from this side. Verified after the update rather than assumed: healthy on backend cupy; /api/jobs went 1 -> 16 against 16 directories on disk, so API and volume agree for the first time; nothing wrongly reclaimed, correct since 16 is under RETAIN=40 and adoption only makes them visible; a real 256^2 plate render completes warm, so the kernel-cache volume survived the image swap. A subsequent render took both counts to 17. The image keeps its explicit [ctk] install and UV_NO_SYNC/UV_OFFLINE pins even though both are now redundant. The header requirement is a property of this slim base, not of the upstream extra, and the cost is measured rather than assumed: uv sync satisfies it first, so the line reports "Audited 1 package" and adds 0.3s to the build. The env pins are now cheap defence-in-depth against any future path that re-enters uv. Docs corrected in place: the README's upstream-finding section is now a resolved-finding record, and the two "bounded ~500 MB" claims say which commit made that bound hold across restarts rather than only within a process. Comment-side changes pushed to the live compose dir; no restart was needed for them.
This commit is contained in:
@@ -9,13 +9,13 @@ irv-ml1:8410. Commits `a2b5b58`, `8189076`.
|
||||
Tracking `main` per operator: PR #4 merged and `main` HEAD was exactly the
|
||||
pinned `8025366`, so tracking-a-moving-ref and keeping-the-pin agreed anyway.
|
||||
|
||||
**`main` has since moved to `464dfc2` (PR #5, 2026-08-19) — the running
|
||||
container is deliberately still on `8025366`.** #5 fixes landmines 1 and 2 at
|
||||
the source (see below). Nothing on the service is broken by staying behind:
|
||||
the image's own `[ctk]` install and env pins already neutralise both, so a
|
||||
rebuild buys reliability that is already present, and the operator has called
|
||||
wind-down on the project. `update.sh` picks up `main` on the next rebuild that
|
||||
has a real reason behind it.
|
||||
**Now deployed at `b72425b` (2026-08-19).** The container sat on `8025366` for
|
||||
a few hours after PR #5 (`464dfc2`) landed — deliberately, since the image's
|
||||
own guards already neutralised both landmines and the project was in
|
||||
wind-down. PR #6 (the job-store rehydrate, operator-green-lit) was the rebuild
|
||||
with a real reason behind it, and one `update.sh` run carried both. Verified
|
||||
end to end after the update: healthy, `backend: cupy`, and a real 256² plate
|
||||
render completes warm — the kernel-cache volume survived the image swap.
|
||||
|
||||
## Build context lives OUTSIDE the compose dir — on purpose
|
||||
|
||||
@@ -74,7 +74,11 @@ in `464dfc2`** (`gpu` is now `cupy-cuda12x[ctk]>=13`). **The explicit install
|
||||
stays in the Dockerfile**: the header requirement is a property of *this*
|
||||
image — a slim base with no system CUDA toolkit — so it belongs in the file
|
||||
that creates the problem, not inherited from an extra two repos away. It also
|
||||
survives any future restructuring of the `gpu` extra.
|
||||
survives any future restructuring of the `gpu` extra. Cost of keeping it is
|
||||
now measured, not assumed: since `uv sync` satisfies it first, the line
|
||||
reports `Audited 1 package in 49ms` and adds **0.3s** to the build. A no-op
|
||||
that documents a non-obvious requirement is worth 0.3s. (waterland-dev
|
||||
independently agreed they would keep it too.)
|
||||
|
||||
## Landmine 3 — the GPU index inside the container is not the host's
|
||||
|
||||
@@ -121,14 +125,34 @@ it is their code. Prune the volume by hand if it bites first.
|
||||
claim holds within one process lifetime and nowhere else, which on a
|
||||
`restart: unless-stopped` service is the wrong lifetime to have bounded. They
|
||||
have **surfaced a startup-rehydrate fix to the operator** rather than opening a
|
||||
third PR during wind-down. **Operator green-lit it (2026-08-19)** — a PR is
|
||||
expected. Volume at **61 MB / 16 job dirs** as of 2026-08-19 08:30Z —
|
||||
unchanged since first observed, so there is no clock on it.
|
||||
third PR during wind-down. **Operator green-lit it; PR #6 merged as `b72425b`
|
||||
and is DEPLOYED (2026-08-19).**
|
||||
|
||||
**This is the rebuild trigger.** When the rehydrate PR merges, one `update.sh`
|
||||
run on irv-ml1 moves the container off `8025366` and picks up `464dfc2`'s
|
||||
fixes in the same motion — the "rebuild with a real reason behind it" named
|
||||
above. Nothing else on the infra side needs coordinating.
|
||||
Startup rehydrate, as recommended — and waterland-dev deliberately went
|
||||
further than the framing I sent them. I had said a directory the scan cannot
|
||||
parse "just does not enter the index"; they made the opposite call, because a
|
||||
directory that never enters the index is exactly the one that never gets
|
||||
reclaimed. **That is the sharper reading and it is the reason the fix works on
|
||||
this volume at all** — the 16 pre-existing dirs have no sidecar. Their
|
||||
adoption ladder: sidecar → restored verbatim; no sidecar → adopted with
|
||||
dimensions recovered from the PNG IHDR (24-byte read, not a decode); corrupt
|
||||
sidecar → degrades to inference, no startup crash; **neither source nor
|
||||
sidecar → skipped on purpose**, since adopting it would turn eviction into a
|
||||
delete-arbitrary-directories primitive pointed at this volume. Sidecar writes
|
||||
go through `os.replace`, and `job.json` is excluded from `ARTIFACTS` so it is
|
||||
unreachable via the artifact route.
|
||||
|
||||
They also closed a second leak I never saw, because it needs a restart
|
||||
*mid-render* to surface: a job left `running`/`queued` in its sidecar is
|
||||
non-terminal forever, and eviction skips non-terminal jobs — so it is a
|
||||
phantom that is never reclaimed and `queue_depth` over-reports for the life of
|
||||
the process. Adoption now marks those `failed`.
|
||||
|
||||
**Verified on this host after the update:** `/api/jobs` went **1 → 16** while
|
||||
the volume stayed at 16 dirs / 61 MB — disk and API agree for the first time.
|
||||
Nothing was reclaimed, correctly: 16 is under `RETAIN=40`, so adoption only
|
||||
made them visible. A subsequent real render took both to 17. From here the
|
||||
store is bounded **across** restarts, not merely within a process.
|
||||
|
||||
## Access
|
||||
|
||||
|
||||
Reference in New Issue
Block a user