079c7b15e3
Three changes prepping infra for asset_engine's orchestration feature
(SSH-driven bring-up / bring-down of irv-ml1 inference services with
per-device VRAM gating, contract in vh/asset-engine commit 5a36f8c):
1. asset-engine compose + .env.example + playbook gain a read-only
bind-mount for /app/runtime/ssh — the dedicated ed25519 keypair
(generated on ana-docker, not in the repo) plus a pinned known_hosts
for irv-ml1's host fingerprint. Env vars SSH_KEY_PATH and
SSH_KNOWN_HOSTS are exposed for the app to consume.
2. docs/asset-engine/services.yaml gains a `lifecycle: { stack, vram_gb,
gpu_device_id }` block on each of 12 orchestratable irv-ml1 services
(kokoro, chatterbox, index-tts, qwen3-tts, cosyvoice, fish-s2,
kyutai-tts, vibevoice, voxtral, parakeet, stable-audio-open, ace-step).
VRAM numbers are estimates from model footprint at fp16 — tune from
real nvidia-smi measurements once the gate is live. comfyui and
kokoro-captioned are deliberately excluded (variable-VRAM and
shared-container respectively).
3. servers/irv-ml1/README.md docker-stacks table now lists all 13
inference stacks (was only dockge + agents + comfyui) with port +
GPU pinning columns.
Pubkey deployed to ~lkraven/.ssh/authorized_keys on irv-ml1;
end-to-end SSH from ana-docker → irv-ml1 verified with strict
host-key checking.
42 lines
1.7 KiB
Bash
42 lines
1.7 KiB
Bash
# asset-engine stack tunables. Copy to `.env` on ana-docker before deploying.
|
|
#
|
|
# The deploy playbook seeds `.env` from this template on first run only —
|
|
# it won't clobber an existing `.env`.
|
|
|
|
# Image tag. Built locally from the asset-engine git repo by the playbook.
|
|
ASSET_ENGINE_IMAGE=asset-engine:local
|
|
|
|
# Host port exposing the FastAPI app (container listens on 8000 internally).
|
|
# Internal tooling, LAN-only — this port is the only entry point. No Traefik.
|
|
ASSET_ENGINE_PORT=8200
|
|
|
|
# Bind address for the host port. 0.0.0.0 = LAN-reachable.
|
|
ASSET_ENGINE_BIND=0.0.0.0
|
|
|
|
# Host paths for state. Container runs as uid 1000 — paths must be writable
|
|
# by that uid (mkdir'd by the playbook without sudo, so lkraven-owned when
|
|
# lkraven is uid 1000 on the host).
|
|
#
|
|
# DB lives separately from outputs so we can grow outputs/ onto a different
|
|
# volume later without restoring DB state on top of it.
|
|
ASSET_ENGINE_DB_DIR=/opt/docker/conf/asset-engine/db
|
|
ASSET_ENGINE_OUTPUTS_DIR=/opt/docker/conf/asset-engine/outputs
|
|
|
|
# SSH key dir for orchestrating irv-ml1 services (bring up / down via SSH +
|
|
# docker compose). Holds id_ed25519 (mode 600) + known_hosts (mode 644)
|
|
# pre-populated with irv-ml1's pinned ed25519 fingerprint. Generated on the
|
|
# host directly so the private key never crosses the network. Bind-mounted
|
|
# read-only into the container at /app/runtime/ssh.
|
|
ASSET_ENGINE_SSH_DIR=/opt/docker/conf/asset-engine/ssh
|
|
|
|
# Inference target. Default is irv-ml1 over WG. Override if the fleet's
|
|
# inference host moves.
|
|
INFERENCE_HOST=10.100.79.3
|
|
|
|
# OIDC seam — empty in v1 (auth is no-op). Populate when v2 forward-auth
|
|
# lands. Pre-allocated here so the surface is visible in the config file
|
|
# before code reads it.
|
|
OIDC_ISSUER=
|
|
OIDC_CLIENT_ID=
|
|
OIDC_CLIENT_SECRET=
|