feat(waterland-studio): deploy b72425b — all three upstream findings fixed

One update.sh run on irv-ml1 carried both open upstream PRs, per the
operator's green-light on the job-store fix:

- #5 (464dfc2) declares cupy-cuda12x[ctk] on the gpu extra and takes uv
  out of the render path (sys.executable -m waterland.cli), retiring the
  runtime prune trap at the source.
- #6 (b72425b) rehydrates the job index from the data volume at startup,
  fixing the unbounded store growth reported from this side.

Verified after the update rather than assumed: healthy on backend cupy;
/api/jobs went 1 -> 16 against 16 directories on disk, so API and volume
agree for the first time; nothing wrongly reclaimed, correct since 16 is
under RETAIN=40 and adoption only makes them visible; a real 256^2 plate
render completes warm, so the kernel-cache volume survived the image swap.
A subsequent render took both counts to 17.

The image keeps its explicit [ctk] install and UV_NO_SYNC/UV_OFFLINE pins
even though both are now redundant. The header requirement is a property
of this slim base, not of the upstream extra, and the cost is measured
rather than assumed: uv sync satisfies it first, so the line reports
"Audited 1 package" and adds 0.3s to the build. The env pins are now
cheap defence-in-depth against any future path that re-enters uv.

Docs corrected in place: the README's upstream-finding section is now a
resolved-finding record, and the two "bounded ~500 MB" claims say which
commit made that bound hold across restarts rather than only within a
process. Comment-side changes pushed to the live compose dir; no restart
was needed for them.
This commit is contained in:
vh
2026-08-19 01:39:13 -07:00
parent a2b6bf409e
commit e6907819b0
4 changed files with 76 additions and 33 deletions
@@ -9,13 +9,13 @@ irv-ml1:8410. Commits `a2b5b58`, `8189076`.
Tracking `main` per operator: PR #4 merged and `main` HEAD was exactly the Tracking `main` per operator: PR #4 merged and `main` HEAD was exactly the
pinned `8025366`, so tracking-a-moving-ref and keeping-the-pin agreed anyway. pinned `8025366`, so tracking-a-moving-ref and keeping-the-pin agreed anyway.
**`main` has since moved to `464dfc2` (PR #5, 2026-08-19) — the running **Now deployed at `b72425b` (2026-08-19).** The container sat on `8025366` for
container is deliberately still on `8025366`.** #5 fixes landmines 1 and 2 at a few hours after PR #5 (`464dfc2`) landed — deliberately, since the image's
the source (see below). Nothing on the service is broken by staying behind: own guards already neutralised both landmines and the project was in
the image's own `[ctk]` install and env pins already neutralise both, so a wind-down. PR #6 (the job-store rehydrate, operator-green-lit) was the rebuild
rebuild buys reliability that is already present, and the operator has called with a real reason behind it, and one `update.sh` run carried both. Verified
wind-down on the project. `update.sh` picks up `main` on the next rebuild that end to end after the update: healthy, `backend: cupy`, and a real 256² plate
has a real reason behind it. render completes warm — the kernel-cache volume survived the image swap.
## Build context lives OUTSIDE the compose dir — on purpose ## Build context lives OUTSIDE the compose dir — on purpose
@@ -74,7 +74,11 @@ in `464dfc2`** (`gpu` is now `cupy-cuda12x[ctk]>=13`). **The explicit install
stays in the Dockerfile**: the header requirement is a property of *this* stays in the Dockerfile**: the header requirement is a property of *this*
image — a slim base with no system CUDA toolkit — so it belongs in the file image — a slim base with no system CUDA toolkit — so it belongs in the file
that creates the problem, not inherited from an extra two repos away. It also that creates the problem, not inherited from an extra two repos away. It also
survives any future restructuring of the `gpu` extra. survives any future restructuring of the `gpu` extra. Cost of keeping it is
now measured, not assumed: since `uv sync` satisfies it first, the line
reports `Audited 1 package in 49ms` and adds **0.3s** to the build. A no-op
that documents a non-obvious requirement is worth 0.3s. (waterland-dev
independently agreed they would keep it too.)
## Landmine 3 — the GPU index inside the container is not the host's ## Landmine 3 — the GPU index inside the container is not the host's
@@ -121,14 +125,34 @@ it is their code. Prune the volume by hand if it bites first.
claim holds within one process lifetime and nowhere else, which on a claim holds within one process lifetime and nowhere else, which on a
`restart: unless-stopped` service is the wrong lifetime to have bounded. They `restart: unless-stopped` service is the wrong lifetime to have bounded. They
have **surfaced a startup-rehydrate fix to the operator** rather than opening a have **surfaced a startup-rehydrate fix to the operator** rather than opening a
third PR during wind-down. **Operator green-lit it (2026-08-19)** — a PR is third PR during wind-down. **Operator green-lit it; PR #6 merged as `b72425b`
expected. Volume at **61 MB / 16 job dirs** as of 2026-08-19 08:30Z — and is DEPLOYED (2026-08-19).**
unchanged since first observed, so there is no clock on it.
**This is the rebuild trigger.** When the rehydrate PR merges, one `update.sh` Startup rehydrate, as recommended — and waterland-dev deliberately went
run on irv-ml1 moves the container off `8025366` and picks up `464dfc2`'s further than the framing I sent them. I had said a directory the scan cannot
fixes in the same motion — the "rebuild with a real reason behind it" named parse "just does not enter the index"; they made the opposite call, because a
above. Nothing else on the infra side needs coordinating. directory that never enters the index is exactly the one that never gets
reclaimed. **That is the sharper reading and it is the reason the fix works on
this volume at all** — the 16 pre-existing dirs have no sidecar. Their
adoption ladder: sidecar → restored verbatim; no sidecar → adopted with
dimensions recovered from the PNG IHDR (24-byte read, not a decode); corrupt
sidecar → degrades to inference, no startup crash; **neither source nor
sidecar → skipped on purpose**, since adopting it would turn eviction into a
delete-arbitrary-directories primitive pointed at this volume. Sidecar writes
go through `os.replace`, and `job.json` is excluded from `ARTIFACTS` so it is
unreachable via the artifact route.
They also closed a second leak I never saw, because it needs a restart
*mid-render* to surface: a job left `running`/`queued` in its sidecar is
non-terminal forever, and eviction skips non-terminal jobs — so it is a
phantom that is never reclaimed and `queue_depth` over-reports for the life of
the process. Adoption now marks those `failed`.
**Verified on this host after the update:** `/api/jobs` went **1 → 16** while
the volume stayed at 16 dirs / 61 MB — disk and API agree for the first time.
Nothing was reclaimed, correctly: 16 is under `RETAIN=40`, so adoption only
made them visible. A subsequent real render took both to 17. From here the
store is bounded **across** restarts, not merely within a process.
## Access ## Access
+1 -1
View File
@@ -113,7 +113,7 @@ _As of 2026-08-19 (late) — a long infra session, all of it landed and committe
- **🟢 FLEET `.internal` DNS — LIVE 2026-08-19.** `<host>.<site>.internal`, sites `ana`/`esh`/`nh3`. `dns/internal.yaml` is the source of truth; `scripts/dns-sync.py` reconciles the three AdGuard resolvers (diff → prompt → apply, idempotent). 42 names resolving from all three sites. **Colo got its first resolver ever** (`stacks/adguard-ana/`, API on **8053** not 8080, no blocklists by design) — before this, ana-docker resolved straight against `1.1.1.1`. Auth = a dedicated `infra-ops` AdGuard user, password vaulted `nh3-dev/adguard-infra-ops-password`. **⏳ TWO OPEN, both operator's to schedule:** (a) colo hosts still point at `1.1.1.1` so they do not yet *use* the new resolver — repointing a site's DNS is a separate change; (b) the static-v6 convention (each server at its site's `/64` with low bits echoing the v4 octet, `esh-docker-vm` → `…::45`) is **proposed, not ruled on**. The `v6:` column is empty and correct — no fleet host has a global v6 address yet. → `persistent-memory.d/2026-08-19-fleet-internal-dns.md` - **🟢 FLEET `.internal` DNS — LIVE 2026-08-19.** `<host>.<site>.internal`, sites `ana`/`esh`/`nh3`. `dns/internal.yaml` is the source of truth; `scripts/dns-sync.py` reconciles the three AdGuard resolvers (diff → prompt → apply, idempotent). 42 names resolving from all three sites. **Colo got its first resolver ever** (`stacks/adguard-ana/`, API on **8053** not 8080, no blocklists by design) — before this, ana-docker resolved straight against `1.1.1.1`. Auth = a dedicated `infra-ops` AdGuard user, password vaulted `nh3-dev/adguard-infra-ops-password`. **⏳ TWO OPEN, both operator's to schedule:** (a) colo hosts still point at `1.1.1.1` so they do not yet *use* the new resolver — repointing a site's DNS is a separate change; (b) the static-v6 convention (each server at its site's `/64` with low bits echoing the v4 octet, `esh-docker-vm` → `…::45`) is **proposed, not ruled on**. The `v6:` column is empty and correct — no fleet host has a global v6 address yet. → `persistent-memory.d/2026-08-19-fleet-internal-dns.md`
- **🟢 WATERLAND STUDIO — containerised + live, irv-ml1:8410.** Displaced a bare `nohup` (PID 1283383). Warm 256²+anim render **6.1s**, beating the 7.4s bare-metal figure. `claude-bot` now has **read** on `vh/waterland` (operator-granted); token on irv-ml1 root-owned 0600, repo-scoped helper, nothing in `.git/config`. **⏳ Reported upstream, NOT patched by us:** the on-disk job store grows without bound across restarts (`RETAIN=40` eviction only sees the in-memory dict) — waterland-dev **confirmed it as a real defect**, and the **operator green-lit their startup-rehydrate fix** — PR expected (61 MB / 16 dirs, no clock on it). Upstream `main` is now `464dfc2` (PR #5 fixes the uv-prune and CUDA-header landmines at source); **the container deliberately stays on `8025366`** and keeps its `[ctk]` install + `UV_NO_SYNC`/`UV_OFFLINE` pins as defence-in-depth. **⏳ NEXT: when the rehydrate PR merges, one `update.sh` on irv-ml1 lands it and `464dfc2` together** — that merge is the rebuild trigger. Thread closed. → `persistent-memory.d/2026-08-19-waterland-studio-containerised.md` - **🟢 WATERLAND STUDIO — containerised + live, irv-ml1:8410.** Displaced a bare `nohup` (PID 1283383). Warm 256²+anim render **6.1s**, beating the 7.4s bare-metal figure. `claude-bot` now has **read** on `vh/waterland` (operator-granted); token on irv-ml1 root-owned 0600, repo-scoped helper, nothing in `.git/config`. **✅ All three upstream findings fixed and DEPLOYED at `b72425b`** (one `update.sh` run, 2026-08-19): PR #5 declared the CUDA headers and took `uv` out of the render path; PR #6 (operator-green-lit) added the startup rehydrate that fixes the unbounded job-store growth. Verified after: `/api/jobs` **1 → 16** against 16 dirs on disk — API and volume agree for the first time — nothing wrongly reclaimed (16 &lt; `RETAIN=40`), healthy on `backend: cupy`, warm render intact. The image **keeps** its `[ctk]` install (+0.3s, now a no-op audit) and `UV_NO_SYNC`/`UV_OFFLINE` as defence-in-depth. Thread closed by both sides. → `persistent-memory.d/2026-08-19-waterland-studio-containerised.md`
- **🟢 HOMEPAGE — cleaned + themed (Australis Skyfall).** Three real defects fixed (UltraSeedbox on all tabs, Uptime Kuma double-rendered, fiction column counts), AI tab reordered by clickability, then themed from the operator's Skyfall handoff bundle with a background generated by **Arbo** (`t2i-ui-background`, job `13f0891f4e42`). ⚠ **After any recreate the tab bar/wallpaper/i18n vanish for up to ~an hour and then return on their own — do not chase it.** ⚠ CSS is served per-request: a theme change needs a **reload**, not a recreate, and candidate CSS can be injected live via Playwright for seconds-long iteration. → `persistent-memory.d/2026-08-19-homepage-skyfall-theme.md` - **🟢 HOMEPAGE — cleaned + themed (Australis Skyfall).** Three real defects fixed (UltraSeedbox on all tabs, Uptime Kuma double-rendered, fiction column counts), AI tab reordered by clickability, then themed from the operator's Skyfall handoff bundle with a background generated by **Arbo** (`t2i-ui-background`, job `13f0891f4e42`). ⚠ **After any recreate the tab bar/wallpaper/i18n vanish for up to ~an hour and then return on their own — do not chase it.** ⚠ CSS is served per-request: a theme change needs a **reload**, not a recreate, and candidate CSS can be injected live via Playwright for seconds-long iteration. → `persistent-memory.d/2026-08-19-homepage-skyfall-theme.md`
+33 -17
View File
@@ -95,8 +95,11 @@ service merely looks slow.
needs wider reach. Ask. needs wider reach. Ask.
- Job store is scratch output, not source-of-truth: ~12 MB per animated job, - Job store is scratch output, not source-of-truth: ~12 MB per animated job,
self-evicting at 40 retained jobs (`RETAIN` in `studio/jobs.py`), so steady self-evicting at 40 retained jobs (`RETAIN` in `studio/jobs.py`), so steady
state is bounded around 500 MB. The 4 jobs from the bare-metal instance were state is bounded around 500 MB. As of `b72425b` the index is rehydrated from
copied in at cutover. disk at startup, so that bound holds **across restarts** and not merely
within one process — see the resolved-finding section below for why that
distinction mattered. The 4 jobs from the bare-metal instance were copied in
at cutover.
- Internal render timeout is 480 s, which is why the healthcheck interval is - Internal render timeout is 480 s, which is why the healthcheck interval is
loose — an aggressive probe would measure queue depth rather than liveness. loose — an aggressive probe would measure queue depth rather than liveness.
@@ -127,21 +130,34 @@ and the grant itself, passed inline and never written to disk on this host.
Do not reintroduce it: a site-admin credential on a GPU box is a blast radius Do not reintroduce it: a site-admin credential on a GPU box is a blast radius
nobody needs for a read-only fetch. nobody needs for a read-only fetch.
## ⚠️ Upstream finding: the on-disk job store grows without bound ## ✅ Resolved upstream: the on-disk job store used to grow without bound
The job index is **memory-only** — `JobStore._jobs` is a plain dict, and **Fixed in `b72425b` and deployed here.** Kept on the record because the
nothing scans `WATERLAND_STUDIO_DATA` at startup. Two consequences: symptom is easy to misread and the fix has a property worth knowing about.
1. After any restart the UI and `/api/jobs` list only jobs created *since* The job index used to be **memory-only** — `JobStore._jobs` was a plain dict
that restart, even though every earlier job's files are still on the volume. and nothing scanned `WATERLAND_STUDIO_DATA` at startup. Two consequences:
This is cosmetic and not data loss — verified at cutover: the API reported
1 job while the volume held all 16 directories (60.6 MB).
2. More importantly, the `RETAIN = 40` self-eviction only ever sees in-memory
jobs, so directories orphaned by a restart are **never** reclaimed. The
handover's "bounded around 500 MB" holds within a single process lifetime;
across restarts the store grows monotonically at ~12 MB per animated job.
Growth is slow, so this is not urgent, but it is unbounded and the volume will 1. After any restart the UI and `/api/jobs` listed only jobs created *since*
need occasional pruning until upstream rehydrates the index (or sweeps orphans) that restart, even though every earlier job's files were still on the
at startup. Reported to waterland-dev — it is their call, not a thing to patch volume. Cosmetic, not data loss — and it is how this was spotted: the API
from the infrastructure side. reported 1 job while the volume held all 16 directories (60.6 MB).
2. The real one: `RETAIN = 40` self-eviction only ever walked the in-memory
dict, so directories orphaned by a restart were **never** reclaimed. The
handover's "bounded around 500 MB" held within a single process lifetime;
across restarts the store grew monotonically at ~12 MB per animated job.
Upstream now **rehydrates the index at startup** by scanning the data dir, and
deliberately *adopts* directories it cannot fully parse rather than skipping
them — a directory that never enters the index is precisely the one eviction
can never reclaim. Sidecar present → restored; no sidecar → adopted with
dimensions read from the PNG IHDR; corrupt sidecar → inferred, no startup
crash. **Neither source nor sidecar → skipped on purpose**, because adopting
those would make eviction a delete-arbitrary-directories primitive pointed at
this volume. Jobs left `running`/`queued` by a mid-render restart are marked
`failed`, so they stop being unreclaimable phantoms that inflate `queue_depth`.
Verified here after the update: `/api/jobs` went **1 → 16** against 16
directories on disk, with nothing reclaimed — correct, since 16 is under
`RETAIN=40`; adoption only made them visible. The bound now holds across
restarts, so no hand-pruning is needed.
+3
View File
@@ -40,6 +40,9 @@ services:
# Job store: uploaded sources plus rendered plates and animations. # Job store: uploaded sources plus rendered plates and animations.
# ~12 MB per job with an animation; the app self-evicts at 40 retained # ~12 MB per job with an animation; the app self-evicts at 40 retained
# jobs (RETAIN in studio/jobs.py), so steady state is bounded ~500 MB. # jobs (RETAIN in studio/jobs.py), so steady state is bounded ~500 MB.
# That bound holds ACROSS restarts only as of b72425b, which rehydrates
# the job index from this volume at startup; before that, eviction was
# blind to anything a restart had orphaned. See the README.
# Scratch output, not source-of-truth — losing it costs a re-render. # Scratch output, not source-of-truth — losing it costs a re-render.
- waterland_studio_data:/data - waterland_studio_data:/data
# cupy JIT kernel cache. Not optional for good behaviour: cupy compiles # cupy JIT kernel cache. Not optional for good behaviour: cupy compiles