Files
esh-pfi-infrastructure/servers/nh3-dev
vh 92a4114b90 feat(nh3-dev): migrate to docker-ce 29.8 + compose plugin; drop compose v1
Operator cleared the swap and ruled out a docker-compose v1 shim. Ran
playbooks/upgrade-docker-ce.yaml: docker.io 20.10.24 -> docker-ce 29.8.0,
docker-compose 1.29.2 -> compose plugin v5.5.1, containerd 1.6.20 ->
containerd.io 2.3.5, buildx v0.37.1 added. 12 changed, 0 failed, verify
4/4. talk and beszel-agent back healthy on their restart policies.

The pre-state was worse than 'old': there was no cli-plugins directory, so
'docker compose' was not a command and exited 0 on a help blurb — a silent
no-op that reads as a successful deploy.

Records two things the run surfaced. vastblue-u5-pg and its anonymous
volume were removed when the old daemon stopped; the playbook has no rm,
prune or purge and five other containers survived, so the cause is almost
certainly --rm, unprovable now that the record is gone. It was measured
beforehand as zero user tables in every database, so nothing was lost. And
the playbook's restart loop runs as infra-ops and cannot read a root-owned
0600 stack .env, so it false-FAILs that stack.

Also notes that nh3-dev is the only host where /opt/docker/compose is
root-owned; the other four are lkraven. Created /opt/docker/compose/talk
as lkraven so tts-dev can move talk out of ~/talk. Normalising the parent
is left to the operator.
2026-09-14 12:46:12 -07:00
..

nh3-dev

NH3-site developer box — 10.100.10.50 (WireGuard-reachable from the NH3 subnet). General-purpose dev VM that hosts agent-fleet sidecars and live Claude Code sessions; not a Docker-stack host in the stacks/ sense.

Reach: ssh 10.100.10.50 (as lkraven), or the dedicated agent identity ssh -i ~/.ssh/infra-ops_ed25519 infra-ops@10.100.10.50 (NOPASSWD sudo). infra-ops bootstrapped here 2026-06-04 (see [reference_infra_ops_sudo_identity] in auto-memory). Note: Claude Code sessions often run natively on this box, so local Bash already executes here — no SSH-to-self needed for non-privileged work.

What runs here

  • NH3 egress proxy — RETIRED 2026-09-06 (replaced by headscale exit nodes; danted disabled, config .retired). Was: durable internal-only SOCKS5 socks5h://10.100.10.50:1080 (dante, ACL'd to the WG net). Residential egress for colo services gated on their datacenter IP (e.g. YouTube bot-gate). Runbook + setup committed; consumers point *_PROXY at it.
  • Docker runtime — docker-ce since 2026-09-14. Was Debian's docker.io 20.10.24 + the Python docker-compose 1.29.2 v1 CLI + containerd 1.6.20, with no cli-plugins directory at all — so docker compose (space) was not a command: it printed a help blurb and exited 0, which a deploy script cannot distinguish from success. Migrated via playbooks/upgrade-docker-ce.yaml to docker-ce 29.8.0 / compose plugin v5.5.1 / containerd.io 2.3.5 / buildx v0.37.1. The v1 docker-compose (hyphen) binary is gone and no shim was installed (operator ruling 2026-09-14) — fix callers, don't paper over them.
    • ⚠ vastblue-u5-pg (an empty postgres:16 probe container, restart: no, plain docker run, anonymous volume) and its volume were removed when the old daemon stopped. The playbook is not the cause — it has no rm, prune or purge, and the other five containers survived, two of them long-exited. Almost certainly --rm / AutoRemove=true, unprovable after the fact because the container record is gone. Measured beforehand as zero user tables in every database, so no data was lost. Lesson: capture AutoRemove and RestartPolicy together when snapshotting a container you are about to bounce.
    • ⚠ The playbook's restart loop runs as infra-ops and cannot read a root-owned 0600 stack .env (/opt/docker/compose/beszel/.env), so it reports that stack as FAILED even when restart: unless-stopped brings it back fine. Playbook-side fix pending.
  • /opt/docker/compose ownership — nh3-dev is the fleet outlier. root:root here; lkraven:lkraven on irv-ml1, nh3-docker, ana-docker and esh-docker-vm. So "can a project session deploy its own stack" is false only on the box where sessions actually run. /opt/docker/compose/talk was created lkraven-owned 2026-09-14 so tts-dev can migrate talk out of ~/talk (a convention violation that hides it from anything walking /opt/docker/compose/*/). Normalising the parent is unresolved — operator's call; /opt/docker itself is a separate three-way split (755 root, 777 root on two hosts, 755 lkraven).
  • ttyd fleet driver-seat — web/iPad seat into the zellij Claude session (ttyd behind Caddy; OSC52 clipboard shim). User systemd services under ~/.config.
  • mead-hall — Bifrost tool-provider sidecar (:5173), CI-deployed from vh/mead-hall.
  • Hermes Agent gateway — OpenAI-compatible agent API on 127.0.0.1:8765 (hermes-gateway.service, user systemd, installed 2026-09-14 via hermes gateway install). Runs the per-user install at ~/.hermes/hermes-agent (v0.21.1, b88e677); config in ~/.hermes/{.env,config.yaml}. Stood up for SVOS/Miranda, which replaced Worldtree with Hermes on 2026-09-11 and cannot boot without it. Bearer auth is mandatory even on loopback — key vaulted as nh3-dev/hermes/api-server-key. ⚠ The gateway registers Hermes's full toolset by default — 28 toolsets, 14 enabled, terminal / code_execution / file / browser among them. SVOS's security model is that write-capable tools are never registered, not that they are refused at dispatch, so platform_toolsets: {api_server: []} is set in config.yaml (2026-09-14) and measured back as 28 rows / 0 enabled / 0 tools on /v1/toolsets. The endpoint still reports all 28 rows with their flags, which is what SVOS's _hermes_roster derives its required-config line from — narrowing does not blind it. Becomes [svos_miranda] once SVOS's plugin lands in $HERMES_HOME/plugins/.
  • Hermes model backend → gen-large on the fleet LiteLLM gateway (free local compute), set 2026-09-14 per operator ruling. Until then model.default said anthropic/claude-opus-4.6 with model.base_url at openrouter, but provider: auto plus a lone zai credential in auth.json silently resolved Miranda to GLM-5.3 on the paid z.ai Coding Plan — three settings that had to be read together before the real answer fell out. Now default: gen-large / provider: custom / base_url: http://10.250.50.70:4000/v1, verified by a real turn (hermes status → gen-large / Custom endpoint, plus a 660-token completion through /v1/chat/completions). The openrouter/nous credit warnings cleared with it. ⚠ CUSTOM_API_KEY / HERMES_CUSTOM_API_KEY are INERT for bare provider: "custom" — they only bind a named custom_providers: entry via its key_env. Set model.api_key in config.yaml instead; its only env fallback is the legacy name OPENROUTER_API_KEY. Get this wrong and the request ships the placeholder no-key-required, LiteLLM 401s inside the response body, and hermes status still reports a perfectly healthy gen-large / Custom endpoint — so status alone cannot verify this change, only a real completion can. The zai credential is still in auth.json, present and unused; provider: custom is explicit so it is not a candidate, and hermes fallback list is empty, so there is no degraded-mode route that quietly re-bills z.ai. Miranda fails rather than fails over if LiteLLM is down.
  • ⚠ Do not rotate nh3-dev/hermes/api-server-key yet. SVOS currently reuses that value as the HS256 signing key on its Bifrost wall, so a rotation would 401 every Miranda tool call. svos-dev is splitting theirs off (operator- approved 2026-09-14) and will confirm when it lands; rotation is safe after that, not before.
  • bloom_music dev — ~/development/bloom_music; its web/ test harness uses Playwright headless Chromium for OSMD browser-geometry assertions.
  • The Booth — ephemeral media drop board (:8090, booth.service), from eshpfi services/booth/. Lets CC sessions surface A/B renders + smoke results (and browser uploads for pickup) to the operator; 24h TTL, Homepage-linked. Since 2026-09-09 it also carries asks — a session poses a multiple-choice question in a booth, the operator answers a radio form + notes in the browser, and the pick lands as an answer sidecar the session reads (booth ask / booth answer --wait). ⚠ The booth CLI is on PATH via ~/.local/bin/booth → services/booth/scripts/booth, symlinked 2026-09-09; before that it was on no PATH at all, so every session following the global link-board convention was hitting command not found unless it used the full path. ~/.zshenv puts ~/.local/bin in PATH for non-interactive ssh nh3-dev '<cmd>' too.
  • jackdaw-compose — JackDAW AI Composer /compose backend (:8787, jackdaw-compose.service), a thin stateless bun server/index.ts from ~/development/jackdaw → LiteLLM gen. Origin-gated (INV-BK04/BK05), reached same-origin via the :4500 bench's /compose proxy. Hosted for jackdaw-dev (their code; the model endpoint + key live in server env only — unit is 0600, not committed).

Box-wide Playwright / Chromium (2026-06-04)

Available to every user/project on this box — no per-home playwright install:

  • System shared-libs: apt-installed via playwright install-deps chromium (Debian-12 set + xvfb), global.
  • Browser binaries: shared /opt/ms-playwright (chromium-1223 + headless-shell + ffmpeg), root-owned, world-readable. Installed via infra-ops.
  • Discovery: PLAYWRIGHT_BROWSERS_PATH=/opt/ms-playwright set globally in /etc/environment (PAM/all sessions) + /etc/profile.d/playwright-browsers.sh (login shells). A project just npm i playwright (skip-browser-download is fine) and resolves the shared binary; verified launching headless from /opt as a normal user.
  • To add more browsers / bump: ssh infra-ops@10.100.10.50 'sudo env PLAYWRIGHT_BROWSERS_PATH=/opt/ms-playwright npx -y playwright install <browser>'.

Notes

  • PFI-owned Linux — in scope for infra-ops management (apt, systemctl, service lifecycle). Added to the fleet bootstrap's Tier 1.
  • OS: Debian 12 (bookworm). See system-details.txt for the latest snapshot (scripts/refresh-server-info.sh nh3-dev).
  • Not in the colo Docker-stack topology — no /opt/docker/compose deploy target; workloads are systemd services + dev checkouts.
  • Retired (2026-06-08): volva.service + heid.service user systemd units removed. Heid/Volva were re-architected from Python systemd daemons (volva run / heid run pollers) into Claude Code session orchestrators (heid commit 12aa5a9); the ~/development/volva dir + venvs are gone. volva.service had been crash-looping 203/EXEC. Cleanup done by infra-ops at heid's request.