gitea-runner: stack + playbook for self-hosted Actions

Central runner on ana-docker (gitea is local; existing fleet tooling
already SSHes from there). Playbook is parameterized so future
site-local runners (nh3-docker, esh-docker-vm) drop in via --var
overrides instead of copy-paste.

Includes a workflow template for vh/task-board that calls the existing
deploy-task-board.yaml playbook — keeps the playbook as the single
source of truth for "how task-board is deployed", manual or automated.

Labels embed `:docker://node:20-bookworm-slim` schema; without it,
act_runner v0.6+ silently falls back to host-mode and runs job steps
inside the Alpine runner container (no apt/python/node), breaking any
real workflow. node:20-bookworm-slim is small + has git + node so
actions/checkout works out of the box.
This commit is contained in:
vh
2026-04-29 18:12:36 -07:00
parent 48aaa53c9d
commit f014d5534a
7 changed files with 517 additions and 2 deletions
+42
View File
@@ -0,0 +1,42 @@
# Gitea instance the runner polls.
GITEA_INSTANCE_URL=https://gitea.phasefinal.com
# Runner identity. NAME shows in the runner list in gitea admin;
# LABELS is what `runs-on:` in workflow YAML matches against.
#
# `pfi-fleet` is the cross-host label used by all PFI runners — keep
# it on every runner. The host-specific label (ana-docker, nh3-docker,
# esh-docker-vm) lets a workflow pin to a particular site if it has to
# (e.g. needs LAN access to a host only reachable from there).
#
# Each label MUST carry a `:docker://<image>` schema. Without it,
# act_runner v0.6+ falls back to "host" mode — jobs would execute
# inside the runner container itself (Alpine, no apt/python/node)
# instead of in a spawned job container, breaking any workflow that
# expects a normal Linux userspace.
#
# `node:20-bookworm-slim` is a sane default: small (~150 MB), has git
# + node (so `actions/checkout@v4` and other JS-based actions work),
# debian-based so `apt-get install` works. Workflows can still override
# with their own job-level `container: image:` if they need something
# different.
GITEA_RUNNER_NAME=ana-docker-runner
GITEA_RUNNER_LABELS=pfi-fleet:docker://node:20-bookworm-slim,ana-docker:docker://node:20-bookworm-slim
# One-time registration token. Generate at:
# https://gitea.phasefinal.com/-/admin/actions/runners (org/global)
# or:
# https://gitea.phasefinal.com/<owner>/<repo>/settings/actions/runners (repo)
#
# Used only on first start; afterwards the runner caches a permanent
# token at ${DATA_DIR}/.runner. Safe to clear after the runner shows
# up in the gitea admin runner list.
GITEA_RUNNER_REGISTRATION_TOKEN=
# Persistence (.runner credentials, build cache, working dirs).
DATA_DIR=/opt/docker/conf/gitea-runner/data
# Pin to a tagged release in production; :latest is fine for
# bootstrapping. Check https://gitea.com/gitea/act_runner/releases
# for the current stable.
RUNNER_IMAGE=gitea/act_runner:latest
+134
View File
@@ -0,0 +1,134 @@
# gitea-runner
Self-hosted [Gitea Actions](https://docs.gitea.com/usage/actions/overview)
runner. Polls `gitea.phasefinal.com` for jobs from any repo that has a
`.gitea/workflows/` directory and runs them in ephemeral docker
containers on this host.
**Server:** ana-docker (single central runner — see "Topology" below
for when to add more)
**Image:** `gitea/act_runner:latest`
**Outbound only** — no host port published; the runner connects out to
gitea, gitea never connects in.
## Topology
We run **one central runner on ana-docker**. Reasoning:
- gitea is on ana-docker, so runner→API is local
- existing fleet tooling (`elway`, `sync-stacks`, `deploy-stack`,
`refresh-server-info`) already SSHes from one origin to all hosts;
the runner inherits that pattern
- single point to manage SSH keys, secrets, and runner upgrades
The deploy playbook (`playbooks/deploy-gitea-runner.yaml`) is
parameterized by host / runner-name / labels, so spinning up
`nh3-docker-runner` or an ESH runner later is a one-line elway
invocation — not a copy-pasted playbook.
**When to add a site-local runner:**
- Cross-site SSH from ana-docker to that site has become unreliable
- A workflow needs LAN access to something only reachable from inside
that site's network segment
- You want failure isolation (NH3 can deploy itself when Anaheim is down)
Until one of those bites, one runner is enough.
## Prereqs
Before running the deploy playbook:
1. **Verify Gitea Actions is enabled.** In gitea 1.21+ Actions ships
on by default, but check `/-/admin/actions` resolves. If not, add
`GITEA__actions__ENABLED=true` to the gitea stack env and bounce.
2. **Generate a registration token.** Pick the scope:
| Scope | URL | Use when |
|---|---|---|
| Global (admin) | `https://gitea.phasefinal.com/-/admin/actions/runners` | runner serves any repo on the instance (recommended for the central PFI runner) |
| Org/user | `https://gitea.phasefinal.com/<owner>/-/actions/runners` | runner serves all repos under one owner |
| Repo | `https://gitea.phasefinal.com/<owner>/<repo>/settings/actions/runners` | runner serves one repo |
Register-as-admin is right for our use case: one runner, fleet-wide.
3. **Create an SSH deploy key for the runner** that lets it execute
the elway playbooks against fleet hosts. The key lives only on
ana-docker (passed in as a workflow secret per repo, or mounted
into the runner via volume — see "Wiring deploys" below).
4. **(Optional) Create a Gitea PAT** with `read:repository` scope on
`vh/esh-pfi-infrastructure`. Workflows need to check out the
management repo to get at the playbooks; the auto-injected
`GITHUB_TOKEN` only works for the triggering repo.
## Deploy
```bash
# Edit .env first if not using defaults — at minimum paste the registration token
$EDITOR stacks/gitea-runner/.env.example # template
# First-time deploy
scripts/elway ana-docker --playbook playbooks/deploy-gitea-runner.yaml
# Site-local runner later (NH3 or ESH)
scripts/elway nh3-docker --playbook playbooks/deploy-gitea-runner.yaml \
--var runner_name=nh3-docker-runner --var runner_labels=pfi-fleet,nh3-docker
```
The playbook seeds `.env` from `.env.example` only if absent; for the
first run, copy `.env.example` to `/opt/docker/compose/gitea-runner/.env`
on the host and paste the registration token in before running, OR
let the playbook seed it and edit on the host before the
`docker compose up -d` step (it's idempotent — second run will pick up
the edited token).
After successful registration, the token is consumed (it's one-time
use). You can clear `GITEA_RUNNER_REGISTRATION_TOKEN` from `.env`;
the runner reads its permanent credentials from `${DATA_DIR}/.runner`
on subsequent starts.
## Wiring deploys
A workflow that runs on the central runner needs three things:
1. **`runs-on:`** matching a runner label — `pfi-fleet` (cross-fleet)
or `ana-docker` (pin to that host).
2. **An SSH key** to reach the deploy target. Stored as a repo or
org-level Actions secret named e.g. `DEPLOY_SSH_KEY`. The
corresponding public key must be in `~lkraven/.ssh/authorized_keys`
on every host the workflow targets.
3. **A token to clone `vh/esh-pfi-infrastructure`** if the workflow
wants to invoke an elway playbook from this repo. Stored as
`MGMT_REPO_TOKEN` (Gitea PAT, `read:repository` scope).
See `stacks/task-board/gitea-workflow-deploy.yaml.example` for a
complete deploy workflow that consumes all three.
## Path layout (on ana-docker)
| Host path | Container path | Purpose | Restic? |
|---|---|---|---|
| `/opt/docker/compose/gitea-runner/` | — | compose.yaml + .env | included (via `/opt/docker`) |
| `/opt/docker/conf/gitea-runner/data/` | `/data` | `.runner` creds, cache, job workspaces | excluded (regenerable; nothing irreplaceable) |
## Operations
```bash
# Tail runner logs
ssh ana-docker docker logs -f gitea-runner
# List currently registered runners (admin)
# https://gitea.phasefinal.com/-/admin/actions/runners
# Re-register (lost the .runner file? regenerate token, then:)
ssh ana-docker docker compose -f /opt/docker/compose/gitea-runner/compose.yaml down
ssh ana-docker rm /opt/docker/conf/gitea-runner/data/.runner
# paste new GITEA_RUNNER_REGISTRATION_TOKEN into .env
scripts/elway ana-docker --playbook playbooks/deploy-gitea-runner.yaml
# Pin to a specific act_runner version
# Edit RUNNER_IMAGE in /opt/docker/compose/gitea-runner/.env, then:
scripts/elway ana-docker --playbook playbooks/deploy-gitea-runner.yaml
```
+38
View File
@@ -0,0 +1,38 @@
# Gitea Actions self-hosted runner.
#
# Polls https://gitea.phasefinal.com for queued jobs and runs them in
# ephemeral job-containers spawned via the host docker socket. The
# runner registers itself on first start using
# GITEA_RUNNER_REGISTRATION_TOKEN; subsequent starts reuse credentials
# cached at ${DATA_DIR}/.runner.
#
# Tunables live in .env. config.yaml controls runner internals (job
# timeouts, capacity, container engine settings) — edit conf/config.yaml,
# don't inline that here.
services:
runner:
image: ${RUNNER_IMAGE}
container_name: gitea-runner
restart: unless-stopped
environment:
- GITEA_INSTANCE_URL=${GITEA_INSTANCE_URL}
- GITEA_RUNNER_REGISTRATION_TOKEN=${GITEA_RUNNER_REGISTRATION_TOKEN}
- GITEA_RUNNER_NAME=${GITEA_RUNNER_NAME}
- GITEA_RUNNER_LABELS=${GITEA_RUNNER_LABELS}
- CONFIG_FILE=/data/config.yaml
volumes:
- ${DATA_DIR}:/data
- /var/run/docker.sock:/var/run/docker.sock
networks:
- tnet
labels:
- homepage.group=Toolchain
- homepage.name=gitea-runner
- homepage.icon=mdi-cog-play
- homepage.description=Gitea Actions self-hosted runner (${GITEA_RUNNER_NAME})
networks:
tnet:
name: traefik-net
external: true
+55
View File
@@ -0,0 +1,55 @@
# act_runner config. Mounted into the runner container at
# /data/config.yaml (CONFIG_FILE env var points here).
#
# Reference: https://docs.gitea.com/usage/actions/act-runner
log:
level: info
runner:
# Persisted registration credentials. Created on first successful
# `register`; reused on subsequent starts.
file: /data/.runner
# Max parallel jobs this runner will accept.
capacity: 2
# Hard timeout per job — covers a wedged docker build, a hung ssh,
# etc. 30m is generous for our deploy workflows (mostly seconds).
timeout: 30m
# Time given to a job to clean up after a SIGTERM before SIGKILL.
shutdown_timeout: 1m
# TLS verification when talking to gitea. KEEP true in prod.
insecure: false
# Polling cadence + per-poll HTTP timeout.
fetch_timeout: 5s
fetch_interval: 2s
cache:
# Provides actions-cache-compatible storage for `actions/cache`.
enabled: true
dir: /data/cache
container:
# Job containers join this docker network. Lets workflow steps
# talk to other compose services (gitea itself, registries, etc.)
# by container name.
network: traefik-net
privileged: false
# Job containers' working dir is mounted under here on the host
# (via the runner's docker.sock spawning). Kept on the runner's
# /data volume so workspaces persist briefly between steps.
workdir_parent: /data/workspace
# Volumes the runner allows job containers to bind-mount. Keep tight.
valid_volumes: []
force_pull: false
host:
workdir_parent: /data/host-workspace