Files
esh-pfi-infrastructure/stacks/asset-engine
vh e0d1c44137 chore(fleet): repoint stale irv-ml1 refs (10.100.79.3 -> irv-ml1.nh3.internal)
The 2026-09-06 headscale cutover retired irv-ml1's wg0 tunnel IP 10.100.79.3
(now 10.6.110.50). Repointed all LIVE canonical refs to the DNS NAME so the next
move can't re-break them: homepage.href/siteMonitor labels across 25 stack
composes, load-bearing env defaults (asset-engine INFERENCE_HOST, open-webui
AUDIO_TTS_OPENAI_API_BASE_URL, skaldsong SKALDSONG_TTS_BASE_URL, zonos-gateway
ZONOS_URL, dia), homepage services.yaml manual cards (Voice Design Studio,
IRV-ML1), and servers/irv-ml1/ssh-target. Updated the stale 'WG tunnel' comment
to the mesh reality.

Left as-is: README curl-examples and .env.example comments (docs), and historical
mentions in CLAUDE.md/persistent-memory. NOTE: applying the label repoints to the
RUNNING irv-ml1 containers needs a recreate per service (labels read at creation);
deployed .env values are separate from these canonical defaults.
2026-09-07 15:08:56 -07:00
..

asset-engine

Control plane over the PFI inference fleet — FastAPI + HTMX/Shoelace UI that exposes the catalog at docs/asset-engine/services.yaml as a web app, routes generation requests to inference hosts (irv-ml1 over WG by default), and persists generated assets to a local SQLite DB + content-addressed blob store.

Server: ana-docker URL: http://10.250.50.70:8200 (configurable via .env) Upstream repo: vh/asset-engine Image: asset-engine:local — built on the host from the git repo by the deploy playbook. Not pulled from a registry.

Deploy

Two paths — automated (preferred) and manual (escape hatch / first-time).

Automated (Gitea Actions, push-to-main)

The asset-engine repo ships .gitea/workflows/{ci,deploy}.yaml. CI runs on PRs (uv sync, pytest, catalog drift check against this repo's docs/asset-engine/services.yaml); the deploy workflow runs on push to main and just calls the elway playbook below pinned to the triggering commit SHA. Drift check is a BLOCKING gate — a PR that vendors a services.yaml mismatched against this repo fails CI and can't merge.

A reference copy of the deploy workflow lives next to this README at gitea-workflow-deploy.yaml.example; the canonical source is in the asset-engine repo. The example header lists the two repo secrets required (DEPLOY_SSH_KEY, MGMT_REPO_TOKEN).

Manual (elway from a workstation)

The playbook owns the full flow: clone/update the source repo, docker build, install compose + seed .env, bring up, verify health.

# First deploy (or update to latest main)
scripts/elway ana-docker --playbook playbooks/deploy-asset-engine.yaml

# Pin to a specific ref (tag, branch, or commit SHA)
scripts/elway ana-docker --playbook playbooks/deploy-asset-engine.yaml --var ref=v0.1.0

Path layout (on ana-docker)

Host path Container path Purpose Restic?
/opt/docker/build/asset-engine/ git checkout used as docker build context excluded
/opt/docker/compose/asset-engine/ compose.yaml + .env included (via /opt/docker)
/opt/docker/conf/asset-engine/db/ /app/runtime/db SQLite (asset_engine.db + WAL) included
/opt/docker/conf/asset-engine/outputs/ /app/runtime/outputs content-addressed blob store included

Network model

Internal tooling, LAN-only. Container port 8000 is published on the host at 0.0.0.0:8200 (configurable via ASSET_ENGINE_BIND / ASSET_ENGINE_PORT); access is direct via http://10.250.50.70:8200. No Traefik, no TLS terminator, no public hostname. If we later need TLS or external access, that's a separate decision.

INFERENCE_HOST defaults to 10.100.79.3 (irv-ml1 over WG). Override in .env if the fleet's inference topology moves.

Catalog drift

docs/asset-engine/services.yaml in this repo is the canonical catalog. The asset-engine repo vendors a copy at data/services.yaml and re-vendors via uv run scripts/sync_catalog.py after upstream changes; CI fails any PR where the vendored copy diverges from this one. Pattern is: edit catalog here → asset-engine re-vendors → both sides commit on the same merge window.

Outputs directory growth

outputs/ grows unbounded in v1 — Asset.retention exists in the schema but the GC sweep isn't wired yet. The plan: Beszel alert when du -sh /opt/docker/conf/asset-engine/outputs crosses ~50 GB, revisit the threshold once we have real growth data. Tracking issue in the asset-engine repo.