Three changes prepping infra for asset_engine's orchestration feature
(SSH-driven bring-up / bring-down of irv-ml1 inference services with
per-device VRAM gating, contract in vh/asset-engine commit 5a36f8c):
1. asset-engine compose + .env.example + playbook gain a read-only
bind-mount for /app/runtime/ssh — the dedicated ed25519 keypair
(generated on ana-docker, not in the repo) plus a pinned known_hosts
for irv-ml1's host fingerprint. Env vars SSH_KEY_PATH and
SSH_KNOWN_HOSTS are exposed for the app to consume.
2. docs/asset-engine/services.yaml gains a `lifecycle: { stack, vram_gb,
gpu_device_id }` block on each of 12 orchestratable irv-ml1 services
(kokoro, chatterbox, index-tts, qwen3-tts, cosyvoice, fish-s2,
kyutai-tts, vibevoice, voxtral, parakeet, stable-audio-open, ace-step).
VRAM numbers are estimates from model footprint at fp16 — tune from
real nvidia-smi measurements once the gate is live. comfyui and
kokoro-captioned are deliberately excluded (variable-VRAM and
shared-container respectively).
3. servers/irv-ml1/README.md docker-stacks table now lists all 13
inference stacks (was only dockge + agents + comfyui) with port +
GPU pinning columns.
Pubkey deployed to ~lkraven/.ssh/authorized_keys on irv-ml1;
end-to-end SSH from ana-docker → irv-ml1 verified with strict
host-key checking.
asset-engine
Control plane over the PFI inference fleet — FastAPI + HTMX/Shoelace
UI that exposes the catalog at docs/asset-engine/services.yaml
as a web app, routes generation requests to inference hosts (irv-ml1
over WG by default), and persists generated assets to a local SQLite
DB + content-addressed blob store.
Server: ana-docker
URL: http://10.250.50.70:8200 (configurable via .env)
Upstream repo: vh/asset-engine
Image: asset-engine:local — built on the host from the git repo by
the deploy playbook. Not pulled from a registry.
Deploy
Two paths — automated (preferred) and manual (escape hatch / first-time).
Automated (Gitea Actions, push-to-main)
The asset-engine repo ships .gitea/workflows/{ci,deploy}.yaml. CI runs
on PRs (uv sync, pytest, catalog drift check against this repo's
docs/asset-engine/services.yaml); the deploy workflow runs on push to
main and just calls the elway playbook below pinned to the triggering
commit SHA. Drift check is a BLOCKING gate — a PR that vendors a
services.yaml mismatched against this repo fails CI and can't merge.
A reference copy of the deploy workflow lives next to this README at
gitea-workflow-deploy.yaml.example;
the canonical source is in the asset-engine repo. The example header
lists the two repo secrets required (DEPLOY_SSH_KEY, MGMT_REPO_TOKEN).
Manual (elway from a workstation)
The playbook owns the full flow: clone/update the source repo,
docker build, install compose + seed .env, bring up, verify health.
# First deploy (or update to latest main)
scripts/elway ana-docker --playbook playbooks/deploy-asset-engine.yaml
# Pin to a specific ref (tag, branch, or commit SHA)
scripts/elway ana-docker --playbook playbooks/deploy-asset-engine.yaml --var ref=v0.1.0
Path layout (on ana-docker)
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
/opt/docker/build/asset-engine/ |
— | git checkout used as docker build context | excluded |
/opt/docker/compose/asset-engine/ |
— | compose.yaml + .env | included (via /opt/docker) |
/opt/docker/conf/asset-engine/db/ |
/app/runtime/db |
SQLite (asset_engine.db + WAL) |
included |
/opt/docker/conf/asset-engine/outputs/ |
/app/runtime/outputs |
content-addressed blob store | included |
Network model
Internal tooling, LAN-only. Container port 8000 is published on the host
at 0.0.0.0:8200 (configurable via ASSET_ENGINE_BIND /
ASSET_ENGINE_PORT); access is direct via http://10.250.50.70:8200.
No Traefik, no TLS terminator, no public hostname. If we later need
TLS or external access, that's a separate decision.
INFERENCE_HOST defaults to 10.100.79.3 (irv-ml1 over WG). Override
in .env if the fleet's inference topology moves.
Catalog drift
docs/asset-engine/services.yaml in this repo is the canonical
catalog. The asset-engine repo vendors a copy at data/services.yaml
and re-vendors via uv run scripts/sync_catalog.py after upstream
changes; CI fails any PR where the vendored copy diverges from this
one. Pattern is: edit catalog here → asset-engine re-vendors → both
sides commit on the same merge window.
Outputs directory growth
outputs/ grows unbounded in v1 — Asset.retention exists in the
schema but the GC sweep isn't wired yet. The plan: Beszel alert when
du -sh /opt/docker/conf/asset-engine/outputs crosses ~50 GB,
revisit the threshold once we have real growth data. Tracking issue
in the asset-engine repo.