Operator freed ComfyUI's VRAM directly, so tts-dev is unblocked and the
window request is withdrawn with comfy-dev.
- Verified it was a model unload, not a stop: comfyui still up 8 days,
same pid, HTTP 200, 18,500 -> 612 MiB. Told comfy-dev explicitly so a
VRAM drop is not misread as a restart of their service.
- The resulting 43.8 GB free is a snapshot, not a floor. ComfyUI is live
and reloads ~18.5 GB on the next render, which puts the real floor at
~25.3 GB against FireRedAudio's ~26 GB requirement. The coordination
shrank from "stop ComfyUI" to "don't render during the bench" rather
than disappearing. Flagged to both; not volunteered on comfy-dev's
behalf.
- Withdrew my caching-allocator hypothesis for the dots-tts VRAM. tts-dev
identified it as their prompt-feature cache, capped at 32 entries on
2026-08-14 after two production incidents. A named mechanism with an
incident history beats a plausible story, and the useful finding is
that 14.43 GB sits inside a cap they deliberately chose.
Read-only probes; nothing on the box was changed.
Memory-only; no version bump per the SemVer SKIP list.
tts-dev asked for an A6000 window for an approved TTS bench and flagged a
3090 VRAM delta. Probed the box and mapped PID to container rather than
taking the reported figures.
- The 18.5 GB process they attributed to the 3090 is comfyui, on the
A6000. And it is 18.5 GB rather than the ~11.8 GB they budgeted, so
stopping it gives ~44.4 GB free, not the tight margin they expected.
- Their "~4 GB unaccounted" on the 3090 is two things: parakeet is a
third tenant the doc figure never counted, and dots-tts alone is
holding 14,430 MiB against a burn-in figure of ~6 GB. The second is
the larger finding and it is theirs to act on; handed over with a
caching-allocator hypothesis and a one-restart discriminating test.
- Restated the GPU ordering foot-gun: device_ids ["1"] is the A6000 in a
container, but a bare native CUDA_VISIBLE_DEVICES=1 gets the 3090.
Window not granted unilaterally — comfyui is comfy-dev's and they are
mid-migration, so the request went to them directly and infra-ops relays.
Ruled that the bench runs as a plain container under lkraven rather than
under /opt/docker/compose/, which is for deployed stacks and would leave
a canonical entry reporting as drift until deleted.
Read-only probes; nothing on the box was changed.
Memory-only; no version bump per the SemVer SKIP list.