Commit Graph
461 Commits
Author SHA1 Message Date
vh 85f8a3a072 fix(phasefinal-web): clean up the multi-page site
The 2026-09-28 expansion grew the one-page design to nine pages by
bolting a nav band under each hero. This makes the cluster one site:
- One site header on every page: the wordmark (home link) and a single
  row of short, consistent labels (Applications, Systems, Hosting, AI,
  Engagement, Track record, Contact). The nav no longer wraps into two
  rows at desktop width or four on a phone.
- A real <h1> on every page (the title was a <p>), a skip link, and one
  footer outside <main> on every page, home included.
- The contact page was unreadable: its section borrowed the home page's
  dark contact-card id, so its copy was grey on near-black (2.3:1) and
  its headings measured 1.0:1. It is an ordinary section now.
- Section numbers only where there is a sequence (home, engagement).
- The spec bar names all four services; it stacks cleanly on a phone.
- Contrast: --text-tertiary 4.3:1 -> 4.96:1, --accent-text 4.36:1 ->
  4.74:1 on paper. Track-record labels drop tracked capitals.
- The grid behind the hero and the contact card is its own layer; the
  contact card loses its left side-tab stripe.
- The stylesheet link is versioned (style.css?v=2026-09-28): CSS is
  cached for an hour, and the new markup must not meet the old styles.
- README: nine pages, the shared structure, the CSS-version rule, and the
  operator's anti-slop waivers (the hero's volt rule, Space Grotesk).

Copy is unchanged apart from nav labels (the brief's legal constraints
hold). Still zero JavaScript and zero external requests. Impeccable
detector (controls 7/0, static + rendered at 1280 and 390, light and
dark): 200 unwaived -> 54, all of them the two waived brand devices;
two identical runs. Operator: "ship the site cleanup, keep the rule and
the font".
2026-09-28 16:29:02 -07:00
vh fd20183cbb feat(blender): extension set live in the GUI; MCP acceptance 9/9, probe made safe-mode compliant 2026-09-28 15:08:22 -07:00
vh 339de3cbcd docs(blender): SurfacePsycho STEP round trip into build123d verified by draupnir 2026-09-28 12:53:36 -07:00
vh 83dc497b40 feat(blender): pinned extension set in a read-only System repo, for the GUI and blender-run --extensions
Blender is now a mandatory stage in draupnir's pipeline (Prime, 2026-09-28), and draupnir asked
for eight add-ons from extensions.blender.org: SurfacePsycho 0.10.4, CAD Sketcher 0.32.1,
3D-Print Toolbox 1.4.1, STEP Importer 1.2.1, Bool Tool 2.1.0, LoopTools 4.7.7, MeasureIt 1.8.4,
3MF Import/Export 2.7.7.

- stacks/blender/extensions.lock pins each by version and archive sha256.
- scripts/blender-extensions sync builds fv-ml1:/tank/blender-extensions/5.2/system with Blender's
  own install-file, pre-warms and byte-compiles it, checks a read-only enable, then swaps it in.
  It refuses while the GUI or a blender-run job holds the old directory.
- conf/scripts/startup/fleet_extensions.py enables every package in the System repo: in a timer
  in the GUI (after the prefs load), and as --python ahead of the caller's args in
  blender-run --extensions (a failed enable exits 1 before the caller's script).
- It also patches SurfacePsycho's sp_overwrite_segment_selection from eval() to literal_eval():
  the eval walked past MCP safe mode (control: unpatched ran code, patched refuses).
- blender-run: --extensions (bind mounts via --mount so a missing source fails instead of being
  created); USER/LOGNAME set, which CAD Sketcher's getpass needs.
- compose.yaml mounts the repo read-only and the hook into the GUI container. NOT yet deployed.
- scripts/blender-probes/extensions_acceptance.py: one operator run per add-on, safe-mode
  compliant. Headless 8/9 online and with --network none; CAD Sketcher sketching is GUI-only.
  A Python audit hook saw no network/process events (positive control fired).
2026-09-28 12:50:57 -07:00
vh 7bd00ae77a feat(phasefinal-web): expand site to a multi-page cluster
Single page grows to nine: custom app development, systems engineering,
hosting, and AI services each get a detail page, plus engagement,
track-record, contact, and privacy. Cross-linked nav and footer on every
page; brand CSS extended in place (still zero external requests, zero JS).

Motivation: an app-store registration review flagged the single-page site
as minimal content. This version is intended to be reverted after approval.
2026-09-28 09:59:43 -07:00
vh d0f68a3b18 feat(blender): blender-run one-shot headless wrapper + FLEETTOOLS entry
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).

Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
2026-09-28 08:39:43 -07:00
vh 179309df3f memory: Worldtree reward config live on demo+personal; Blender MCP registration is per task (Prime) 2026-09-27 14:39:43 -07:00
vh ac1cd29afa feat(blender): agent control via mcp-for-blender (in-container, ssh stdio)
The MCP server (mcp-for-blender 2.1.1, frozen requirements) runs inside the Blender
container. Its add-on is vendored at upstream 41a18432 (MIT) and started by a
startup hook. scripts/blender-mcp carries the stdio over ssh + docker exec, so the
add-on socket, which runs arbitrary Python with no auth, stays on the container's
localhost with no published port. It also runs there because viewport screenshots
need a filesystem shared by server and Blender. Telemetry is off and safe mode is
on. The hook also defaults Cycles to OptiX on GPU 3, because safe mode forbids
agents from touching preferences.

Verified end to end from nh3-dev: 36 tools; a GPU render of an agent-built scene;
a viewport screenshot; and safe mode refusing 'import os'. Blender left down
(on demand).
2026-09-27 14:34:07 -07:00
vh 7ae7193c21 feat(blender): Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand (Prime)
Uses linuxserver/blender (Selkies Wayland desktop, NVENC), digest-pinned. The
container sees only GPU 3, has restart "no", and runs only while in use,
because GPU 3 is the reserve card for a full-size vLLM seat. Web desktop on
:3001 with basic auth (vault fv-ml1/blender-web-password). /work is on /tank
and is not backed up; /config lives under /opt/docker (restic).

Acceptance: Cycles finds the card on OptiX and CUDA (sm_120 kernels ship in the
build). The heavy self-test renders in 3.54 s on OptiX vs 22.71 s on CPU, a
functional check with n=1. Web auth answers 401 without credentials and with a
wrong password, and 200 with the right one.
2026-09-27 13:57:35 -07:00
vh 0b8632a7ed fix(esh-docker-vm): /32 route to Home Assistant over macvlan-shim
The shim holds 10.0.50.47/24, which gives two equal connected 10.0.50.0/24 routes,
and ens18's wins. Host-to-HA traffic therefore left via the macvlan parent and was
dropped. HA lost MQTT to the broker on this host on 2026-08-19, 2026-09-21 and
2026-09-25 (the last lasted two days). This adds an ifupdown if-up.d hook that
routes 10.0.50.46/32 via macvlan-shim; /etc/network/interfaces is not edited.
Verified: the route resolves via the shim, the host pings HA, HA reaches :1883,
and HA reconnected to the broker. Diagnosis by ha-dev.

Also corrects the zigbee2mqtt acceptance note, which had wrongly reported HA as
connected.
2026-09-27 12:50:05 -07:00
vh c7b32418e1 feat(zigbee2mqtt): Zigbee2MQTT 2.14.1 on esh-docker-vm, replacing HA's ZHA
Requested by ha-dev; approved by Prime in this session. Radio: SLZB-MR1U chip 0
(EFR32MG21, EmberZNet 8.0.2) at tcp://10.0.90.10:6638, adapter ember. A fresh
network was formed on channel 25, PAN 0xCFF4. The broker is reached as mosquitto
user zigbee2mqtt, with HA discovery on homeassistant/. The frontend on :8099 is
token-protected.

State and the network key stay host-only in /opt/docker/data/zigbee2mqtt
(root 0700, restic). The repo carries only compose, .env.example and the README.
configuration.yaml refers to the secrets as !secret.yaml, and those references
survived Z2M's v4->v5 settings migration.
2026-09-27 12:43:49 -07:00
vh 47cad33dd1 fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object
state's last value ends in ')', ';' or '}', the JSON that follows re-merges two
tokens back, so score_shared refused the request with 422. The engine now wraps
semif_phase1.shared._state_prefix to keep only the tokens the full prompts
share. Each row scores the same token sequence; only the prefill/suffix split
moves.

Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
folded). It checks that the hook is callable and is what score_shared resolves,
that an ordinary state keeps upstream's whole prefix, and that a merge-prone
state scores through the shared path.

Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or
authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72;
the miss is a bf16 tie that flipped across a plain restart (see README).
2026-09-27 10:23:14 -07:00
vh 7e11cf247b spike(semif): SemIf as Cicada's mood source is slower and less apt (no service change)
Against talk /face's guided pose (first paragraph 246 ms median), SemIf in
parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms).
Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM.
Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through
mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%).

README: rotations cost options^2 in suffix tokens, and /decide/shared returns
422 when an object state's last value ends in ) ; or }.
2026-09-27 09:47:28 -07:00
vh 89301a89fe docs(litellm): correct scalar-judge passthrough access semantics
The comment claimed 'gateway-key-gated'; v1.97.0 actually gates auth=true
passthrough routes on per-key metadata allowed_passthrough_routes (OSS path,
not Enterprise), 401 unauthenticated. Grants applied live via /key/update for
all-agents-local and the worldtree gateway key; comment-only change here,
picks up with the next conf deploy.
2026-09-27 09:29:50 -07:00
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00
vh 069725c4b3 feat(semif): SemIf option-logit decisions on fv-ml1 GPU 1 (Prime)
services/semif-serve is a FastAPI wrapper around SemIf's direct and shared torch
scorers (SemIf-OpenJev @ 23cf1f39, MIT). Upstream ships only a batch CLI. The
wrapper loads the pinned Qwen3.5-4B (851bf6e8, BF16) once from the offline HF
cache and returns SemIf's result dicts unchanged, with an optional per-workload
temperature-calibrated view. Contract: semif-serve.contract.md. Built with a
short contract, TDD (39 tests, fake engine and fake torch, no GPU) and a heid
bug-hunt panel (pending).

On the card:
- torch 2.10.0+cu128 with sm_120 kernels, which is SemIf's own stack;
- a hard 12 GiB VRAM cap.
Two defects surfaced only on the card, and each fix is covered by a test:
- 0.1.1: an OOM raised as a chained exception kept the failed request's tensors
  alive (11.9 GiB after the 503). It is now raised unchained, after gc.
- 0.1.2: a large request left 12.6 GB reserved on the shared card. After each
  call, reserved memory over the baseline + 512 MiB is now released.

Acceptance against SemIf's committed torch predictions (authored144):
- 142/144 same top choice; both misses are exact bf16 ties;
- 144/144 identical prompt hashes;
- deterministic A-vs-A;
- negative control 14/144;
- shared vs direct 72/72.
21 binary criteria over one state take 159 ms. The shared-mode capacity table
under the cap is in stacks/semif/README.md.

The Dockerfile installs dependencies from a manifest with the project version
blanked, so a version bump reuses the ~4 GB torch layer. Verified: 41 s rebuild,
dependency layer CACHED.

DNS: semif.fv.internal. Token: vault semif/api-token.
2026-09-27 02:36:56 -07:00
vh 30f2c977b7 fix(restic): vm-esh-nas off env-file too; infra-ops now provisioned there
Prime bootstrapped infra-ops on vm-esh-nas with playbooks/bootstrap-infra-ops-user.yaml.
It got the fleet-pinned uid/gid 850, NOPASSWD sudo with log_output, the docker
group and a 0700 home. That let playbooks/restic-repository-file.yaml migrate
the last restic host: the live profile matched the repo's pre-change sha,
its units no longer carry the URL, its secrets are vaulted, and the live and repo
profiles now match (a5ea75ea). All eight restic hosts are clean.

The staged helper script is gone, both from Prime's home on the host and from the
repo. The docs that described vm-esh-nas as lkraven-only are updated.
2026-09-27 01:56:04 -07:00
vh 6e203dcb99 fix(restic): stop publishing rest-server passwords in systemd units
resticprofile schedule copies env-file values into the generated units, which
are 0644, so RESTIC_REPOSITORY (the rest-server basic-auth password included)
was readable by every local user on every restic host.

New playbooks/restic-repository-file.yaml:
- derives /etc/restic/repository (root 0400) from restic.env;
- uploads the profile switched to repository-file, but only when the live
  profile's sha matches the repo copy it was edited from (drift guard);
- checks the repository is reachable through the new profile (cat config);
- regenerates the units and verifies they exist and contain no rest:http.

Applied to ana-docker, fv-ml1 (configs/restic/ana-ml2), esh-docker-vm,
esh-vm-db, irv-ml1, nh3-dev and nh3-docker. An independent check across all
eight restic hosts (these seven plus esh-ml1) found 0 leaking units. nh3-docker's
scheduled unit ran a real backup afterwards (snapshot a29b889d). Each host's
URL and passphrase are vaulted as <host>/etc/restic/{repository,password}.

restic.env is kept (root 0600) because the per-host READMEs and the freshness
probe source it. A rotation must update the vault, restic.env and repository.

vm-esh-nas has no infra-ops account. Its in-place migration script is staged
for Prime to run with sudo, and its repo profile is pre-edited to match.

Also mirrors augaman-dev's 401df2d (compose header only). The config hash on
esh-ml1 is unchanged.
2026-09-27 01:51:05 -07:00
vh c698751bee sync(restic): pull live esh-docker-vm + irv-ml1 profiles into the repo
Both hosts had live improvements the repo never recorded:
- esh-docker-vm excludes ESPHome's 539 MB of PlatformIO cache (2026-09-14).
- irv-ml1 runs /etc/restic/arbo-checkpoint.sh before the backup, a non-fatal
  SQLite online-backup of arbo's gallery DB. That script is added here too.
The live copies were correct and are the source for the next change. (Also fixes
the fv-ml1 augaman removal time to ~0130 PT.)
2026-09-27 01:38:36 -07:00
vh ad484c3c99 chore(augaman): remove the fv-ml1 instance (Prime)
Prime removed the second instance after the v0.1.3 bench. esh-ml1 handles a face
in ~48 ms, sits in the house next to the cameras, and holds the verified
backup. fv-ml1's gallery was empty (0 identities). The container, gallery volume,
image, compose dir (with its .env), backup dir and build sources are removed from
fv-ml1. GPU_ID / CARD_SUFFIX stay in the compose for any future second host.
2026-09-27 01:31:25 -07:00
vh ef64a69e30 feat(augaman): v0.1.3 on esh-ml1 and fv-ml1; after-bench: GPU ~3x faster, CPU mode regressed
Both hosts are rebuilt from tag v0.1.3 (one ONNX session per detector canvas) and
redeployed. pytest -m gpu tests/vision passes 3/3 on each card.

Server-side for one face, same harness as the v0.1.2 baseline:
- esh-ml1 GPU 144 -> 48 ms
- fv-ml1 GPU 75 -> 27 ms
End to end from nh3-dev: 101.6 and 73.8 ms.

CPU mode got slower on every CPU target: fv-ml1 cpuset 0-5 went 152 -> 205 ms
with a face, and the no-face frame roughly doubled. That is well outside the
run-to-run spread. The suspected cause (not measured) is per-session ORT
thread pools spinning. Reported to augaman-dev. Neither deployment uses CPU
mode.

On esh-ml1 the dependency layer missed the build cache and the rootfs touched
90% until the v0.1.2 image was removed. fv-ml1's build hit the cache, and the
exported requirements are identical, so the stack README now says to check
disk before building on esh-ml1.
2026-09-27 01:13:01 -07:00
vh 317868dc7e feat(augaman): second, fixtures-only instance on fv-ml1 GPU 1; CPU vs GPU speed bench (v0.1.2 baseline)
Prime asked for augaman on fv-ml1's utility card, beside vllm-coder. Mirror
augaman-dev's f77164f compose, which parameterises the GPU reservation (GPU_ID,
default 0) and the Homepage card name (CARD_SUFFIX). esh-ml1's resolved config is
unchanged: same config hash, no recreate.

On fv-ml1: augaman:0.1.2 built on-box from the tag, GPU_ID=1, healthy on CUDA
at 1264 MiB, and pytest -m gpu tests/vision passes 3/3 on the Blackwell. It has
its own gallery and no gallery backup, so it is fixtures-only. The host's raw
restic copy of /var/lib/docker/volumes is not a consistent SQLite backup.

docs/pfi/augaman-speed-bench/ holds the harness (augaman-dev's recipe plus a
no-face control frame and a face-count check on every response), the raw rows
and the summary. Server-side, one face:
- esh-ml1 GPU 144 ms
- fv-ml1 GPU 75 ms
- fv-ml1 CPU on 6 cores 152 ms
- esh-ml1 CPU 888 ms
It agrees with augaman-dev's independent esh-ml1 measurement once each
harness's floor is subtracted. This is the before for v0.1.3's detector fix.
2026-09-27 00:28:20 -07:00
vh d8f59a15d9 feat(esh-ml1): restic backup of augaman's gallery; augaman v0.1.2
esh-ml1 is outside vzdump, so augaman's face gallery reaches backup only
through restic. New playbooks/esh-ml1-restic.yaml installs restic 0.14.0 (the
same Debian package as the other ESH hosts) and resticprofile 0.33.1 (pinned,
sha256-checked). It uploads configs/restic/esh-ml1/ and schedules a daily
0100 PT backup plus a Sunday 0500 PT check to rest-server-ana. The CT runs UTC,
so both schedules name the zone explicitly.

pre-backup.sh is fail-closed: it runs augaman's own backup CLI, and any failure,
including a stopped container, aborts the run. Tested with a stub docker that
exits 1: the run returned 1, and neither the snapshot count nor last-success
moved. The restore was verified at identity level against augaman-dev's
public-domain canary (snapshot fd3061a1: the restored copy's digest over
identities and samples matches the live gallery). That meets the operator gate
for real enrollments.

The repository URL is read through repository-file rather than restic.env.
resticprofile schedule copies env-file values into world-readable systemd
units, which publishes the rest-server password on the env-file hosts
(observed on esh-docker-vm). This is recorded in the backups runbook under
Known gaps, and the playbook verifies no generated unit contains the URL.

esh-ml1 is added to the freshness check's expected ana-side repos and to the
runbook tables.

augaman moves to v0.1.2 (dependency layer keyed on the lock without the
project; per-crop embedding). pytest -m gpu tests/vision passes 3/3 on the
card, and the canary survived the container recreate.
2026-09-27 00:06:33 -07:00
vh c26f7c94e4 feat(augaman): deploy v0.1.1 on esh-ml1:8040 (face recognition for Cicada)
Mirror pfi/augaman deploy/compose.yaml as stacks/augaman, with an .env.example and
a README carrying the biometric backup gate. The image is built on esh-ml1 from a
git archive of the release tag, because the box holds no gitea credentials.

Serving on CUDA and visible in nvidia-smi. The gallery backup is not wired yet
(esh-ml1 has no restic), so only public-domain fixtures may be enrolled.
The on-box gpu test fails its batch-vs-single tolerance 3/3; reported to
augaman-dev, who owns the contract.
2026-09-26 23:50:44 -07:00
vh 2fdbac63d5 feat(vibevoice-asr-seat): switch to Q8_0 (Prime); WER 2/69 vs 3/69 on the bundled clips, +1.1 GB VRAM 2026-09-26 16:16:32 -07:00
vh 8e7ae0675d feat(esh-matter): Matter server (matter.js 1.4.0) on a VLAN-90-only LXC for Home Assistant
For ha-dev (operator-approved 2026-09-26). CT 111 on esh-pve at 10.0.90.20:
Matter/Thread IPv6 (Echo ULA + RA route-information) is link-only, so the
server sits on esh-iot and HA reaches it over routed IPv4 ws :5580.
- playbooks/esh-matter-lxc.yaml: kernel RA (accept_ra=1,
  rt_info_max_plen=64), forwarding off, Docker ip-forward/iptables off;
  nftables admits 5580 from HA 10.0.50.46 only and SSH from mgmt ranges;
  the CT is added to esh-pve's vzdump job (fabric credentials).
- stacks/matter-server: ghcr.io/matter-js/matterjs-server:1.4.0 (digest),
  host networking, /data on the CT.
- Acceptance: fdad:: SLAAC, ping6 thermostat, 2 Thread RIO routes learned, ws
  server_info from inside the HA container; 5580 refused from 10.0.50.45,
  nh3-dev and a temporary VLAN 90 netns vantage.
2026-09-26 13:07:22 -07:00
vh 944bb36031 docs(lfm-vl-uncensored-seat): passed brokkr's 6-image eval; seat stays 2026-09-26 00:58:48 -07:00
vh 9cc3824aec feat(nh3-ml1): abliterated LFM2.5-VL-3B parallel seat :8032 for brokkr's NSFW-caption A/B (direct only) 2026-09-26 00:57:04 -07:00
vh 50c85e0e8a feat(litellm): load-share embed/rerank across esh-ml1 + nh3-ml1 with intra-group failover
- qwen3-embedding and reranker: second deployment on nh3-ml1 (config);
  reranker-a3-bge-v2-m3: second DB deployment via /model/new.
- router_settings.enable_weighted_failover: true. Only the three
  multi-deployment groups are affected.
- Measured by stopping nh3-ml1 TEI: without failover 7/40 embeds 500'd;
  with it rerank 80/80, embed 38/40 at onset and 60/60 over a 34 s outage;
  nh3 rejoins rotation after restart. Split 12/28 embed, 22/18 rerank.
2026-09-26 00:54:23 -07:00
vh f2792183d4 feat(nh3-ml1): LFM2.5-VL-3B (llama.cpp) + VibeVoice-ASR-Streaming-1.5B (audio.cpp) utility seats
For brokkr's dataset foundry (operator-approved 2026-09-26, relayed).
- stacks/lfm-vl-seat: llama.cpp server-cuda b11176 (digest-pinned), Q5_K_M +
  mmproj Q8_0, :8030; gateway alias lfm25-vl-3b (LiteLLM restarted, 36 s).
  Positive control exact; null control shows it describes a missing image.
- stacks/vibevoice-asr-seat: audio.cpp v0.8.2-audio8-perf-hotfix (the GGUF's
  own runtime, not vibevoice.cpp) on cuda 12.8 runtime + libgomp + libsoxr,
  sha256-pinned; :8031 direct. LibriSpeech WER 3/69, RTF 0.07-0.14; ~31 s
  cold first request.
2026-09-26 00:41:01 -07:00
vh 45484a0007 docs(flash-next-seat): correct gateway alias list against live /model/info (summarizer/classifier are gen-small's) 2026-09-26 00:07:50 -07:00
vh e4cab7f430 docs(flash-next-seat): orcarouter 2026-09-18 update is a V100 repack script only; weights identical, nothing to adopt 2026-09-26 00:02:11 -07:00
vh ddd67df2e9 revert(coder-seat): keep Qwen2.5-Coder on fv-ml1; nh3-ml1 copy removed (5x slower, same quality) 2026-09-25 23:58:08 -07:00
vh 7f066a4b79 feat(coder-seat): Qwen2.5-Coder-1.5B copy on nh3-ml1 (not yet in the gateway)
Same model revision (df3ce67, blob sha matches fv-ml1), vLLM v0.24.0 and flags;
0.33 of the Ada (~5.4 GB, fv-ml1's absolute budget). Healthy, KV 79,808 tokens.

Parity (teacher-forced true code, 40 FIM prompts, 2,215 tokens, cache-salted):
each host is bit-exact with itself; cross-host |dlogprob| median 0.049, top-1
agreement 0.966 (Blackwell vs Ada + fp8 KV); against ground truth no quality
difference (mean logprob diff +0.008 +/- 0.019 SE). Speed on-box: ~5x slower
(64-tok FIM p50 ~1.0 s vs ~0.2 s; 63 vs 338 tok/s). Gateway coder-fast left on
fv-ml1 pending Prime's call.
2026-09-25 23:56:34 -07:00
vh b3b75c16f4 feat(nh3-pve): move AMT to nh3-mgmt (UDM port 6 native VLAN 250) + Homepage link
- PFI-UDMSE port 6 override: native nh3-mgmt, tagged VLANs blocked (was
  forward all / native default). Reservation nh3-pve-amt -> 10.100.250.61.
- nh3-pve: arp_ignore=8 / arp_announce=2 on enp88s0 (now on vmbr0's untagged
  L2) so the host never answers ARP for 10.100.250.60 with the AMT port's MAC.
- DNS nh3-pve-amt.nh3.internal -> 10.100.250.61.
- Homepage: NH3-PVE-AMT card under Infra - NH3 (no siteMonitor/ping: AMT drops
  ICMP and its legacy-renegotiation TLS fails Homepage's fetch).
- AMT keeps its old 10.100.0.151 lease until rebind/expiry (~1920-2224 PT
  2026-09-26); it does not re-DHCP on a VLAN change or link drop (measured).
2026-09-25 22:59:42 -07:00
vh 6d901c0e8d fix(beszel): GPU-LXC temperature alerts watch the GPU, not the hypervisor CPU
nh3-ml1's first Temperature alert (2026-09-25 2105, 88.8 C) was nh3-pve's CPU
package during the nightly vzdump; the GPU sat at 47 C. hwmon is not namespaced,
so an LXC's agent reports its host's coretemp sensors under the LXC's name.

- hosts/{nh3,esh}-ml1.yaml: SENSORS=-coretemp_*,acpitz (blacklist); GPU and
  NVMe temperatures verified still reported.
- Hub: Temperature >95 C / 5 m on nh3-pve and esh-pve for the CPUs (i9-13900H
  TjMax 100 C; nh3-pve 2-hour averages reach 88 C on busy days).
- nh3-pve README: runs warmer than esh-pve; workload confounds it, check
  airflow on the next visit.
2026-09-25 21:13:08 -07:00
vh 5960526c3f feat(nh3-ml1): second TEI embed/rerank backend live on nh3-pve; parity-verified vs esh-ml1
NH3 site visit done: Secure Boot off, iGPU restored as boot VGA, AMT port cabled.

- nh3-pve: NVIDIA 580.178.04 (DKMS, open modules) via pve-nvidia-host.yaml.
- nh3-ml1 = CT 109 @ 10.100.50.80 via gpu-lxc.yaml; embed-rerank (TEI 1.9.4)
  deployed with HOST_NAME/HOST_IP labels.
- Parity vs esh-ml1 (1,126 texts, 2 runs/host, controls): embed cosine min
  0.999993 = own noise floor; overlap@10 1.000 vs MRL-256 positive control
  0.684; rerank top-1 1.00, max diff 0.0014 vs floor 0.0020. On-box speed
  identical within rep spread.
- gpu-lxc.yaml: first step upgrades lxc-pve to >= 6.0.0-2 (Proxmox fix #7006).
  With 6.0.0-1 every docker run in a nesting CT failed on runc 1.5's sysctl
  reopen; applied on nh3-pve (one package).
- pve-nvidia-host.yaml: document that the headers meta drags in the newest
  kernel (nh3-pve went 6.8.12-11 -> -43 at the next reboot).
- Monitoring: Beszel NVIDIA agent + 5 alerts, Kuma #29/#30, Homepage
  nh3-ml1-docker, Dozzle agent (hub 8 clients). DNS nh3-ml1.nh3.internal.
- nh3-pve README: SB/IGFX/driver/kernel state, btmtk oops on -4x kernels,
  AMT cabled but unreachable on the network.

Gateway routing to nh3-ml1 is not changed.
2026-09-25 15:55:33 -07:00
vh bc278d4ba8 refactor(playbooks): host-generic GPU host + GPU LXC playbooks for nh3-ml1
- esh-pve-nvidia-host -> pve-nvidia-host: headers/dkms/build-essential step,
  nouveau blacklist + guarded unload (refuses if nouveau bound a device)
- esh-ml1-lxc -> gpu-lxc: host vars have no defaults (elway aborts on undefined),
  rootfs storage/startup order parameterized, CT kept out of all-guests vzdump jobs
- embed-rerank: Homepage labels take HOST_NAME/HOST_IP, defaults = esh-ml1
2026-09-25 14:17:31 -07:00
vh 2118449881 feat(nh3-pve): prepare for GPU install — pin NIC names by MAC, pull AMT port from vmbr0
nh3-pve and esh-pve are the same Minisforum MS-01 (BIOS AHWSA.1.17). With a
card in the x16 slot its root port takes bus 01 and every NIC moves down a
bus (measured on esh-pve), so predictable names change (enp2s0f0np0 ->
enp3s0f0np0 etc.) and vmbr0 would boot with no uplink. systemd .link files
now pin all NICs by MAC, baked into every initramfs and synced to the ESP;
udev confirms the files apply. The AMT-capable I226-LM (enp88s0) leaves
vmbr0's bridge-ports in the file (next boot), so cabling it for AMT cannot
loop the STP-less bridge.

Also: documented the NanoKVM (https://10.100.250.171) as nh3-pve's console
OOB and that AMT is not wired; nh3-dev's Beszel agent no longer binds NAS
shares (it died on the last NH3 cold start); post-boot checklist in
persistent-memory.
2026-09-25 10:56:17 -07:00
vh 422a27cc1d chore(homepage): commit the live 'hermes-gateway seat' card rename (operator handle split, 2026-09-24) 2026-09-25 09:20:36 -07:00
vh 65dc586497 feat(esh-ml1): wire telemetry and monitoring; fix Dozzle's stale agent list
- Beszel: NVIDIA agent (stacks/beszel hosts/esh-ml1.yaml), hub system
  registered with Status/Disk/CPU/Memory/Temperature alerts; GPU samples
  verified (RTX 2000E Ada util, VRAM, power, temp).
- Uptime Kuma: health monitors for the TEI embed (:8001) and rerank (:8013)
  services — the sole backends behind the gateway, so the one exception to
  "seats are out of Kuma's lane". Reward seat excluded (no consumer).
- Homepage: dockerd on tcp/2375 bound to 10.0.50.80 (playbook step);
  esh-ml1-docker added to docker.yaml.
- Dozzle: agent v10.4.1 on esh-ml1. The hub's DOZZLE_REMOTE_AGENT still
  named ana-ml2's 10.250.50.54 and irv-ml1's retired 10.100.79.3; repointed
  to 10.251.50.54 / 10.6.110.50 (fv-ml1 and irv-ml1 logs were missing).
- Documented nh3-dev's Beszel agent failing on NH3 cold start (bind of an
  automounted NAS share); revived by hand, fix still open.
2026-09-25 09:20:36 -07:00
vh 52612cbe96 feat(reward-seat): move Skywork reward seat from fv-ml1 to esh-ml1; audit finds nothing superseding it
- Audit: Skywork-Reward-V2-Llama-3.1-8B is still #1 of 188 on AllenAI's
  RewardBench 2 per-sample results; no Skywork V3; the -40M sibling is
  vendor-marked experimental. Our AWQ W4A16 quant: 0.847 vs published bf16
  0.860 on a 150-prompt sample (within +/-2.9 pt SE), 96.2% pairwise
  agreement. Double BOS from vLLM on pre-templated text costs a further
  ~2.7 pts; callers must send add_special_tokens=false.
- No working consumer: 0 requests since 2026-09-13; Worldtree Domari points
  at a dead IP with a non-vLLM schema (reported to worldtree-dev).
- Move: sha256-identical model copy; vLLM v0.24.0 on esh-ml1 :8003 at 0.55
  util (KV 1.30x of a 16k request). Parity vs fv-ml1: 149/150 verdicts,
  99.8% pairwise signs, raw |delta| median 0.049.
- Gateway /scalar-judge passthrough -> 10.0.50.80:8003; fv-ml1 vllm-reward
  removed (~10.2 GB freed on GPU 1). stacks/vllm now holds only vllm-coder.
2026-09-25 09:09:48 -07:00
vh 7bdac80878 feat(embed-rerank): TEI is the fleet embed/rerank engine; esh-ml1 sole backend; retire fv-ml1 seats
Prime, 2026-09-25: TEI serves embedding + reranking for esh-ml1 and the fleet
from now on; fv-ml1 retires both once esh-ml1 is up.

- stacks/embed-rerank: vLLM -> TEI 1.9.4 (89- Ada build), same ports
  8001/8013, fail-closed truncation (--auto-truncate false; embed
  --max-batch-tokens 32768).
- litellm: qwen3-embedding -> esh-ml1 only (hosted_vllm/, unchanged address);
  reranker -> huggingface/ provider at :8013 (hosted_vllm/ 422s on TEI's
  `texts` body). DB alias reranker-a3-bge-v2-m3 patched to the same target.
- Verified via the gateway against the retiring fv-ml1 seats: embed cosine
  median 0.999927 (n=203); rerank top-1/top-3 29/30.
- stacks/vllm: vllm-embed and vllm-rerank-a3 removed (containers retired on
  fv-ml1, GPU 1 freed ~6.1 GB); reward + coder unchanged.
- Bake-off record moved to docs/pfi/embed-rerank-tei-vs-vllm-bakeoff.md;
  CLAUDE.md gains the TEI convention.
2026-09-25 08:30:53 -07:00
vh 49b4bf0177 fix(tei-bakeoff): fail-closed truncation; record Dvalin review and nevermore threshold check 2026-09-25 08:17:40 -07:00
vh 582b150131 feat(tei-bakeoff): TEI 1.9.4 vs vLLM on esh-ml1 — parity holds, not faster, much lighter 2026-09-25 08:05:17 -07:00
vh 5402568b76 feat(esh-ml1): RTX 2000E Ada on esh-pve serves embed + rerank as a LiteLLM failover
esh-pve: NVIDIA 580.178.04 (open modules, DKMS) installed on the host from
NVIDIA's .run and loaded live, no reboot. nvidia-persistenced unit creates the
device nodes before pve-guests; the T400's vfio-pci ids and `blacklist nvidia`
retired. playbooks/esh-pve-nvidia-host.yaml.

esh-ml1: CT 110, unprivileged Debian 12, 10.0.50.80, GPU nodes via devN,
NVIDIA userspace from the same .run (--no-kernel-modules), docker-ce +
nvidia-container-toolkit (no-cgroups). playbooks/esh-ml1-lxc.yaml. Not in the
vzdump job on purpose. DNS esh-ml1.esh.internal.

stacks/embed-rerank: Qwen3-Embedding-0.6B :8001 + bge-reranker-v2-m3 :8013 on
the same vLLM v0.24.0 digest and flags as fv-ml1. Measured parity: embed
cosine FV-vs-ESH median 0.999908 (min 0.999772), inside both self-noise
floors; rerank max |delta| 0.000145 vs floor 0.000181, identical ranking.

litellm: qwen3-embedding and reranker gain an esh-ml1 deployment at order 2
behind fv-ml1 (order 1). Order fallback proven with throwaway groups:
refused primary +0.15 s, host-down primary ~18.7 s per call, dead-only 500.

Also: repaired the DB-only alias reranker-a3-bge-v2-m3, dead since the
fv-ml1 relocation (still named 10.250.50.54); documented the third
unkillable homepage wedge on esh-docker-vm.
2026-09-24 22:38:55 -07:00
vh e6da607767 chore(task-board): mothball it; superseded by the High Seat and ledger
Operator ruling 2026-09-24. On ana-docker the stack is `docker compose
down`: the container is removed and port 7878 is closed. Kept for revival:
- the data dir /opt/docker/conf/task-board/data (tasks.db, last written
  2026-09-11)
- the task-board:local image
- stacks/task-board/ and the host's compose + .env

The Uptime Kuma monitor (id 3) was deleted before the stop so it could
not page, and its row is removed from monitors.yaml. Homepage drops the
card on its own, since it reads the container's labels.

Hooks: the container log showed no hook POSTs in 30 days. The only
traffic was open browser tabs holding /events, and the plugin was already
uninstalled on nh3-dev. Removed the paragraph that told sessions to call
task_* tools (CLAUDE.md, and the fork-fleet.sh template that seeds new
repos), the vestigial TASK_BOARD_SESSION env in .claude/settings.json, and
the listings in README and FLEETTOOLS.
2026-09-24 09:22:50 -07:00
vh 5178fdea3c fix(homepage): The High Seat icon -> mdi-eye-outline
svos-dev's call and the better one: The High Seat is the English name for
Hlidskjalf, the seat Odin watches all the worlds from, and watching every
session at once is what the board does. mdi-monitor-dashboard described the
artifact; the eye describes the job.

Deployed and verified in /api/services, not assumed.
2026-09-22 08:33:05 -07:00
vh 0ef25dfdb6 feat(homepage): add The High Seat (SVOS board, nh3-dev:8770)
Requested by svos-dev relaying the operator, 2026-09-22. Reversible work, so
the relay is fine to act on without escalating.

Manual services.yaml entry rather than container labels, because SVOS is a
user-level systemd unit (svos.service) on nh3-dev and nh3-dev is NOT one of
the five hosts in docker.yaml -- Homepage has no Docker API to discover it
through. Same reason the Booth, WhereTF, talk and the infra-hermes seat are
listed by hand, and the comment says so at the entry.

⚠ siteMonitor is "/" deliberately. There is no /api/health on this service:
that path 404s, and a monitor pointed at it would report the board
permanently down while it serves perfectly. svos-dev flagged it and it is
verified here -- / returns 200 and serves the SPA (<title>The High Seat</title>).

Group is Apps, which exists in settings.yaml's layout with tab: Main. An
invented group name gets no tab and renders on ALL tabs, which is how
Scriberr's "AI Systems" leaked across the whole dashboard in August.

Icon mdi-monitor-dashboard is my choice -- svos-dev explicitly did not guess
at one and offered to take a different suggestion.

Verified in a browser, not just in the API: the card renders in Apps with a
green site-monitor at 28 ms.
2026-09-22 08:30:58 -07:00
vh 94899d6fa3 feat(uptimekuma): normalize names off Homepage, publish the status page, restore the widget
NAMES. Homepage already answers "what is this service called", so the monitor
name is now that name verbatim -- a second naming authority is how drift starts,
and an alert reading "[Uptime Kuma] Beszel hub is DOWN" sends you hunting for a
card that does not exist. Only two rows moved (Beszel hub -> Beszel, Dozzle hub
-> Dozzle); the " hub" suffixes were mine, not the services'.

The remaining mixed case is deliberate and is now documented as such. talk, vor
and task-board are lowercase on Homepage and in their own repos; title-casing
them here would make this board disagree with both. What actually looked messy
was scripts/kuma's own ASCII-ordinal sort, which buried every lowercase name
below every capitalised one. Fixed to case-insensitive.

⚠ RENAME SAFETY, which this pass needed and did not have. The seed keys on NAME,
so editing a name would have read as a brand-new monitor: added fresh, with the
old row orphaned, still checking, still alerting, and holding all the history.
`rename_from:` names the old row for one run. Verified: both renamed monitors
kept their IDs and all 67 heartbeats.

Added with it, an orphan warning for any row on the board the spec no longer
names -- because a forgotten monitor keeps paging. Its first cut diffed against
the PRE-EDIT snapshot and so cried wolf on its own successful renames; it
re-reads the board now. A warning that fires on its own correct work is worse
than no warning.

STATUS PAGE + WIDGET. The Homepage uptimekuma widget reads a PUBLISHED status
page (/api/status-page/<slug>), not the admin API -- which is why the widget
labels were deliberately absent from the rebuild: a dashboard widget pointed at
a 404 is the suspected mechanism behind both of Homepage's unkillable D-state
wedges, so shipping one on purpose would have been daft.

The page now exists at slug `nethealth` (the pre-rebuild slug, so old references
still resolve) and is DECLARED IN monitors.yaml, applied by `kuma seed`. Same
principle as the notification channel: a from-scratch rebuild restores the page,
the channel and the monitors together, and nothing the widget depends on lives
only in Kuma's database. Verified in a browser: "13 SITES UP / 0 SITES DOWN /
100% UPTIME" on the dashboard.

⚠ saveStatusPage calls imgDataUrl.startsWith() unconditionally, so passing null
throws and leaves the page CREATED BUT EMPTY -- which reads as success from
/api/status-page (200, correct title) while the group list is silently blank.
Pass "" instead. Commented at the call site.
2026-09-22 00:55:29 -07:00