Files
vh 66c860d6c1 fix(sweep): retire the dead 10.100.79.3 address across the fleet
Operator-directed. The wg0 lifeline retired at the 2026-09-06 headscale cutover is
on no interface anywhere, so anything pointing at it gets no route at all. Homepage
went from 9 dead cards to 0 of 112.

The load-bearing part is that there is no single right target: it depends on who
resolves it. The operator's browser and the Homepage and open-webui containers on
esh-docker-vm all resolve nh3.internal, so those get the name and survive the next
renumber. Containers on irv-ml1 and ana-docker cannot resolve it at all, so those
get the IP.

litellm on ana-docker looked like a counterexample and is not: it resolves the name
only through its own extra_hosts entry, while asset-engine on the same host fails on
it. Test from the container you are about to change, never from a neighbour. Before
committing to the name I confirmed the Homepage container actually fetches ytvc's
healthz through it in production rather than assuming resolution implies reach.

On irv-ml1, 24 files swept and 14 comment-only hits left as port-allocation history.
Seven running containers recreated so the labels took. Seven dormant ones carried
stale labels because editing a compose file does not touch an existing container
object - fixed with compose create --force-recreate, which rebuilds the container
without starting it, the right tool for a deliberately dormant stack.

The sweep's real find was off irv-ml1 entirely: four live values on two other hosts,
silently dead for nine days and alerting nobody. Open WebUI's read-aloud TTS,
asset-engine's inference host, and two skaldsong TTS URLs. Both running services were
recreated and verified reaching their targets afterwards rather than merely carrying
the new string.

One self-inflicted outage worth recording: I recreated breeze-tts for a cosmetic
label change and took ext-tts down for its ~90s CUDA-graph warm-up, returning 500. I
caught it only because I had taken a baseline before touching it. A label-only edit
still costs a full model reload on a GPU container.
2026-09-15 08:41:44 -07:00

42 lines
1.7 KiB
Bash

# asset-engine stack tunables. Copy to `.env` on ana-docker before deploying.
#
# The deploy playbook seeds `.env` from this template on first run only —
# it won't clobber an existing `.env`.
# Image tag. Built locally from the asset-engine git repo by the playbook.
ASSET_ENGINE_IMAGE=asset-engine:local
# Host port exposing the FastAPI app (container listens on 8000 internally).
# Internal tooling, LAN-only — this port is the only entry point. No Traefik.
ASSET_ENGINE_PORT=8200
# Bind address for the host port. 0.0.0.0 = LAN-reachable.
ASSET_ENGINE_BIND=0.0.0.0
# Host paths for state. Container runs as uid 1000 — paths must be writable
# by that uid (mkdir'd by the playbook without sudo, so lkraven-owned when
# lkraven is uid 1000 on the host).
#
# DB lives separately from outputs so we can grow outputs/ onto a different
# volume later without restoring DB state on top of it.
ASSET_ENGINE_DB_DIR=/opt/docker/conf/asset-engine/db
ASSET_ENGINE_OUTPUTS_DIR=/opt/docker/conf/asset-engine/outputs
# SSH key dir for orchestrating irv-ml1 services (bring up / down via SSH +
# docker compose). Holds id_ed25519 (mode 600) + known_hosts (mode 644)
# pre-populated with irv-ml1's pinned ed25519 fingerprint. Generated on the
# host directly so the private key never crosses the network. Bind-mounted
# read-only into the container at /app/runtime/ssh.
ASSET_ENGINE_SSH_DIR=/opt/docker/conf/asset-engine/ssh
# Inference target. Default is irv-ml1 over WG. Override if the fleet's
# inference host moves.
INFERENCE_HOST=10.6.110.50
# OIDC seam — empty in v1 (auth is no-op). Populate when v2 forward-auth
# lands. Pre-allocated here so the surface is visible in the config file
# before code reads it.
OIDC_ISSUER=
OIDC_CLIENT_ID=
OIDC_CLIENT_SECRET=