Operator ruling: root:docker, not a personal username and not a new admin account. lkraven is one of three names he uses, so baking it into shared infrastructure guarantees a stale owner later; a dedicated deploy account buys nothing the existing docker group doesn't, since that group already exists on every host holding exactly lkraven + infra-ops. Applied to nh3-dev, nh3-docker, esh-docker-vm, irv-ml1, ana-docker. All five now 2775 root:docker on /opt/docker and /opt/docker/compose. Clears the 0777 on nh3-docker and ana-docker. 55 stack .env files normalized to root:docker 0640, tightening 43 world-readable ones and opening 31 that were readable by only one of the two deploy identities. No containers bounced — inode metadata only, and .env is read at compose up. Deliberately not a recursive chmod. Three acme.json files and an ssh private key are mode 0600 and traefik/ssh refuse to start if that widens, which would have been a delayed failure surfacing at the next restart rather than now. Protection is both mode-based (0600/0400 untouched) and name-based (acme.json, *.key, *.pem, *.pfx, id_*); modes are symbolic so the 53 executable files in these trees keep their exec bit. Two defects found and fixed mid-rollout. The name list was initially reported but not enforced, so a .key already at 0644 on esh-docker-vm was widened to 0664 — reverted, and the list is now enforced in the chgrp and widening steps. And the exec-bit verify asserted every .sh is executable, which was never true and false-FAILED irv-ml1; it now compares the executable-file count against a recorded baseline.
11 KiB
11 KiB
nh3-dev
NH3-site developer box — 10.100.10.50 (WireGuard-reachable from the NH3
subnet). General-purpose dev VM that hosts agent-fleet sidecars and live Claude
Code sessions; not a Docker-stack host in the stacks/ sense.
Reach: ssh 10.100.10.50 (as lkraven), or the dedicated agent identity
ssh -i ~/.ssh/infra-ops_ed25519 infra-ops@10.100.10.50 (NOPASSWD sudo).
infra-ops bootstrapped here 2026-06-04 (see [reference_infra_ops_sudo_identity]
in auto-memory). Note: Claude Code sessions often run natively on this box, so
local Bash already executes here — no SSH-to-self needed for non-privileged work.
What runs here
- NH3 egress proxy — RETIRED 2026-09-06 (replaced by headscale exit nodes;
danteddisabled, config.retired). Was: durable internal-only SOCKS5socks5h://10.100.10.50:1080(dante, ACL'd to the WG net). Residential egress for colo services gated on their datacenter IP (e.g. YouTube bot-gate). Runbook + setup committed; consumers point*_PROXYat it. - Docker runtime —
docker-cesince 2026-09-14. Was Debian'sdocker.io20.10.24 + the Pythondocker-compose1.29.2 v1 CLI +containerd1.6.20, with nocli-pluginsdirectory at all — sodocker compose(space) was not a command: it printed a help blurb and exited 0, which a deploy script cannot distinguish from success. Migrated viaplaybooks/upgrade-docker-ce.yamlto docker-ce 29.8.0 / compose plugin v5.5.1 / containerd.io 2.3.5 / buildx v0.37.1. The v1docker-compose(hyphen) binary is gone and no shim was installed (operator ruling 2026-09-14) — fix callers, don't paper over them.- ⚠
vastblue-u5-pg(an emptypostgres:16probe container,restart: no, plaindocker run, anonymous volume) and its volume were removed when the old daemon stopped. The playbook is not the cause — it has norm,pruneorpurge, and the other five containers survived, two of them long-exited. Almost certainly--rm/AutoRemove=true, unprovable after the fact because the container record is gone. Measured beforehand as zero user tables in every database, so no data was lost. Lesson: captureAutoRemoveandRestartPolicytogether when snapshotting a container you are about to bounce. - The playbook's restart loop runs as the deploy identity, not root, and a
stack
.envmay legitimately be root-owned0600(/opt/docker/compose/beszel/.env) — compose bails before doing anything, so the stack reported FAILED whilerestart: unless-stoppedhad already brought it back healthy. Fixed 2026-09-14: the loop retries undersudo -nbefore calling it a failure. Verified against beszel — plainrc=1 permission denied, sudo retryrc=0 Container beszel-agent Started. A false FAILED in automation output is worse than a quiet one; it trains readers to skim the failure lines.
- ⚠
/opt/dockerownership — normalised fleet-wide 2026-09-14 toroot:docker 2775(setgid) viaplaybooks/normalize-docker-tree.yaml, operator ruling. Was a three-way split:root:root 755here,root:root 777on nh3-docker + ana-docker (world-writable, from a 2024chmod -R 777to get a git clone working),lkraven 755elsewhere. Not a personal username (lkravenis one of three the operator uses) and not a new admin account — thedockergroup already existed on every host holding exactlylkraven+infra-ops. ⚠ This is not privilege separation:dockermembership is root-equivalent. A future non-root deployer needs a dedicateddeploygroup. Stack.envfiles went toroot:docker 0640— previously 31 of 74 were0600readable by only ONE of the two deploy identities (varying by file, which is what false-FAILED beszel during the docker-ce upgrade) and 43 were world-readable0644. No containers were bounced; these are inode metadata changes and.envis read only atcompose up./opt/docker/compose/talkwas createdlkraven-owned 2026-09-14 and tts-dev migratedtalkinto it the same day — it had been at~/talk, a convention violation that hid it from anything walking/opt/docker/compose/*/. Old path parked at~/talk.migrated-20260914;version: "3.8"dropped from its compose file (it existed only to satisfy the v1 CLI, which is now gone). Normalising the parent directory is unresolved — operator's call;/opt/dockeritself is a separate three-way split (755root,777root on two hosts,755lkraven).- ttyd fleet driver-seat — web/iPad seat into the zellij
Claudesession (ttyd behind Caddy; OSC52 clipboard shim). User systemd services under~/.config. - mead-hall — Bifrost tool-provider sidecar (
:5173), CI-deployed fromvh/mead-hall. - Hermes Agent gateway — OpenAI-compatible agent API on
127.0.0.1:8765(hermes-gateway.service, user systemd, installed 2026-09-14 viahermes gateway install). Runs the per-user install at~/.hermes/hermes-agent(v0.21.1,b88e677); config in~/.hermes/{.env,config.yaml}. Stood up for SVOS/Miranda, which replaced Worldtree with Hermes on 2026-09-11 and cannot boot without it. Bearer auth is mandatory even on loopback — key vaulted asnh3-dev/hermes/api-server-key. ⚠ The gateway registers Hermes's full toolset by default — 28 toolsets, 14 enabled,terminal/code_execution/file/browseramong them. SVOS's security model is that write-capable tools are never registered, not that they are refused at dispatch, soplatform_toolsets: {api_server: []}is set inconfig.yaml(2026-09-14) and measured back as 28 rows / 0 enabled / 0 tools on/v1/toolsets. The endpoint still reports all 28 rows with their flags, which is what SVOS's_hermes_rosterderives its required-config line from — narrowing does not blind it. Becomes[svos_miranda]once SVOS's plugin lands in$HERMES_HOME/plugins/. - Hermes model backend →
gen-largeon the fleet LiteLLM gateway (free local compute), set 2026-09-14 per operator ruling. Until thenmodel.defaultsaidanthropic/claude-opus-4.6withmodel.base_urlat openrouter, butprovider: autoplus a lonezaicredential inauth.jsonsilently resolved Miranda to GLM-5.3 on the paid z.ai Coding Plan — three settings that had to be read together before the real answer fell out. Nowdefault: gen-large/provider: custom/base_url: http://10.250.50.70:4000/v1, verified by a real turn (hermes status→gen-large/Custom endpoint, plus a 660-token completion through/v1/chat/completions). The openrouter/nous credit warnings cleared with it. ⚠CUSTOM_API_KEY/HERMES_CUSTOM_API_KEYare INERT for bareprovider: "custom"— they only bind a namedcustom_providers:entry via itskey_env. Setmodel.api_keyinconfig.yamlinstead; its only env fallback is the legacy nameOPENROUTER_API_KEY. Get this wrong and the request ships the placeholderno-key-required, LiteLLM 401s inside the response body, andhermes statusstill reports a perfectly healthygen-large/Custom endpoint— so status alone cannot verify this change, only a real completion can. Thezaicredential is still inauth.json, present and unused;provider: customis explicit so it is not a candidate, andhermes fallback listis empty, so there is no degraded-mode route that quietly re-bills z.ai. Miranda fails rather than fails over if LiteLLM is down. - ⚠ Do not rotate
nh3-dev/hermes/api-server-keyyet. SVOS currently reuses that value as the HS256 signing key on its Bifrost wall, so a rotation would 401 every Miranda tool call. svos-dev is splitting theirs off (operator- approved 2026-09-14) and will confirm when it lands; rotation is safe after that, not before. - bloom_music dev —
~/development/bloom_music; itsweb/test harness uses Playwright headless Chromium for OSMD browser-geometry assertions. - The Booth — ephemeral media drop board (
:8090,booth.service), from eshpfiservices/booth/. Lets CC sessions surface A/B renders + smoke results (and browser uploads for pickup) to the operator; 24h TTL, Homepage-linked. Since 2026-09-09 it also carries asks — a session poses a multiple-choice question in a booth, the operator answers a radio form + notes in the browser, and the pick lands as an answer sidecar the session reads (booth ask/booth answer --wait). ⚠ TheboothCLI is on PATH via~/.local/bin/booth→services/booth/scripts/booth, symlinked 2026-09-09; before that it was on no PATH at all, so every session following the global link-board convention was hittingcommand not foundunless it used the full path.~/.zshenvputs~/.local/binin PATH for non-interactivessh nh3-dev '<cmd>'too. - jackdaw-compose — JackDAW AI Composer
/composebackend (:8787,jackdaw-compose.service), a thin statelessbun server/index.tsfrom~/development/jackdaw→ LiteLLMgen. Origin-gated (INV-BK04/BK05), reached same-origin via the:4500bench's/composeproxy. Hosted for jackdaw-dev (their code; the model endpoint + key live in server env only — unit is0600, not committed).
Box-wide Playwright / Chromium (2026-06-04)
Available to every user/project on this box — no per-home playwright install:
- System shared-libs: apt-installed via
playwright install-deps chromium(Debian-12 set + xvfb), global. - Browser binaries: shared
/opt/ms-playwright(chromium-1223 + headless-shell + ffmpeg), root-owned, world-readable. Installed viainfra-ops. - Discovery:
PLAYWRIGHT_BROWSERS_PATH=/opt/ms-playwrightset globally in/etc/environment(PAM/all sessions) +/etc/profile.d/playwright-browsers.sh(login shells). A project justnpm i playwright(skip-browser-download is fine) and resolves the shared binary; verified launching headless from/optas a normal user. - To add more browsers / bump:
ssh infra-ops@10.100.10.50 'sudo env PLAYWRIGHT_BROWSERS_PATH=/opt/ms-playwright npx -y playwright install <browser>'.
Notes
- PFI-owned Linux — in scope for
infra-opsmanagement (apt, systemctl, service lifecycle). Added to the fleet bootstrap's Tier 1. - OS: Debian 12 (bookworm). See
system-details.txtfor the latest snapshot (scripts/refresh-server-info.sh nh3-dev). - Not in the colo Docker-stack topology — no
/opt/docker/composedeploy target; workloads are systemd services + dev checkouts. - Retired (2026-06-08):
volva.service+heid.serviceuser systemd units removed. Heid/Volva were re-architected from Python systemd daemons (volva run/heid runpollers) into Claude Code session orchestrators (heid commit12aa5a9); the~/development/volvadir + venvs are gone.volva.servicehad been crash-looping203/EXEC. Cleanup done by infra-ops at heid's request.