Commit Graph
1434 Commits
Author SHA1 Message Date
vh 5bbf0aaeba feat(intern-decision-serve): Intern-Decision-4B behind semif-serve's HTTP surface
Contract, service and tests (fake engine, no GPU). Scores through the checkpoint's own
inference.py (DecisionEngine.predict, sha256-pinned); maps semif decisions onto Jev choice
questions, packs /decide/shared into calls of at most 16, runs orderings in waves, and keeps
semif's error mapping, admission, body limit and hard VRAM cap. Deltas from semif-serve are
listed in the contract.
2026-09-30 09:04:39 -07:00
vh bb806e3596 memory: U11b step-5 auto-trigger (3 PASS) is mine; semif->intern-decision in flight; scriberr GPU budget 2026-09-30 09:03:21 -07:00
vh 0176ec0a5b scriberr: fit GPU 1 beside intern-decision — 120 s Parakeet slices + expandable_segments
Measured Parakeet peak on a 35-min file (n=3 each, deterministic): 300 s
9,384 MiB, 120 s 6,510, 60 s 5,976, 10 s 5,634 (fixed floor); with
expandable_segments 120 s 5,496 and 60 s 5,502. Verified 5,496 under the
recreated container's own env. Transcripts: 95.8% word-sequence similarity
vs 300 s, diffs mostly casing/punctuation.
2026-09-30 09:02:45 -07:00
vh d9bbaa07b2 memory: Jev bench done — Intern-Decision-4B is the SemIf replacement candidate 2026-09-30 05:01:30 -07:00
vh 475d6d6bcb docs: Jev candidate bench vs SemIf (fv-ml1 GPU 3) — Intern-Decision-4B is the replacement if SemIf is displaced
Operator ask relayed by brokkr-smithy-dev. Positive control (SemIf 187/231,
hard 0.613) reproduced exactly; negative control and a 4-restart noise floor
(0 flips) measured. On our replaced-baseline sets no candidate beats SemIf-with-
rotations beyond the ~4-pt floor; Intern-Decision-4B native matches it at one
ordering, is better on Wyrd, fits 9.7/10.3 GB and is 1.5-2.3x faster. JevBench
rank does not transfer. Raw per-item data kept out of git.
2026-09-30 05:00:45 -07:00
vh 9a6ac59da7 memory: U11b legacy-memory archive done (dedicated restic repo, drill passed); destroy-by 2026-10-30 runbook 2026-09-30 03:01:54 -07:00
vh 37d0b34682 memory: U11 daily off batches live (infra-hermes), first PASS; U11b /embed usage + legacy inventory + archive plan 2026-09-30 02:24:28 -07:00
vh aee8de3395 wt-memory-gate-batch: busy is exit 4 (retry later), and PASS is reported too
Both verdicts now build the n for the U11b data-deletion gate (three
consecutive PASS batches at off). First real run 20260930T090608Z via the
detached-worktree path: PASS, user median 0.83.
2026-09-30 02:24:05 -07:00
vh 6dd9ed2964 memory: U11a overnight log sweep clean but traffic-free; re-sweep TODO 2026-09-30 01:55:01 -07:00
vh 21e064b588 memory: scriberr 35-min retry succeeded with SemIf offline 2026-09-30 01:39:07 -07:00
vh 9daf43683d semif: offline by operator ruling — scriberr needs GPU 1 headroom
Stopped (not removed) 2026-09-30 0135 PT. Scriberr's Parakeet path hardcodes
5-minute slices; attention memory is quadratic in slice length, so a long file
needs >6 GB and hit CUDA OOM with SemIf resident (~6.7 GB free). GPU 1 now
81,806 MiB used. Durable fix (shorter scriberr slices) deferred to later.
2026-09-30 01:36:18 -07:00
vh 5186b3dc9a scripts: wt-memory-gate-batch — one Worldtree U8 gate batch against demo's deployed sha
For infra-hermes's daily batch through the U11 off window (operator ruling
2026-09-29 2340; mode revised to off 2026-09-30). Enforces worldtree-dev's
terms: flock plus a refusal while any harness batch runs, code under test =
demo's deployed sha (a detached worktree when the main tree has moved; refuses
on pyproject/uv.lock/packages drift the shared venv cannot honour), no git
commit, no retry. Exit 0 PASS / 1 FAIL / 2 harness error / 3 refused.
2026-09-30 01:26:46 -07:00
vh 060fd92e4e memory: Worldtree U11a — personal flipped to legacy off and verified 2026-09-30 01:21:13 -07:00
vh 9905b76404 memory: Worldtree U11a — Prime ruled legacy off; demo flipped via config repo and verified; personal pending 2026-09-30 01:17:23 -07:00
vh f9fae53599 memory: snapshot — Worldtree U10 done, U11a prepped, U8 wrapper for infra-hermes in progress; Blender extensions live; Bonsai spike closed; 3 entries archived 2026-09-29 23:42:03 -07:00
vh 274b8175c0 memory: Worldtree U11a prepped (staged config on demo, U8 window harness facts) 2026-09-29 23:33:24 -07:00
vh 47d8fd6139 memory: Worldtree U10 backfill committed on personal (797 filed) 2026-09-29 10:07:45 -07:00
vh 40427b4822 memory: U10 personal presync; demo U9 forget-policy gap fixed; pinned chroma in container layer 2026-09-29 08:55:16 -07:00
vh b7b6e01d6e memory: Worldtree U10 backfill committed on demo; memory_tagger config sync needed on personal 2026-09-28 23:52:05 -07:00
vh ea5d8dfd4e docs(blender): desktop Blender + MCP registration per working session, not per task (Prime 2026-09-28) 2026-09-28 16:36:44 -07:00
vh 85f8a3a072 fix(phasefinal-web): clean up the multi-page site
The 2026-09-28 expansion grew the one-page design to nine pages by
bolting a nav band under each hero. This makes the cluster one site:
- One site header on every page: the wordmark (home link) and a single
  row of short, consistent labels (Applications, Systems, Hosting, AI,
  Engagement, Track record, Contact). The nav no longer wraps into two
  rows at desktop width or four on a phone.
- A real <h1> on every page (the title was a <p>), a skip link, and one
  footer outside <main> on every page, home included.
- The contact page was unreadable: its section borrowed the home page's
  dark contact-card id, so its copy was grey on near-black (2.3:1) and
  its headings measured 1.0:1. It is an ordinary section now.
- Section numbers only where there is a sequence (home, engagement).
- The spec bar names all four services; it stacks cleanly on a phone.
- Contrast: --text-tertiary 4.3:1 -> 4.96:1, --accent-text 4.36:1 ->
  4.74:1 on paper. Track-record labels drop tracked capitals.
- The grid behind the hero and the contact card is its own layer; the
  contact card loses its left side-tab stripe.
- The stylesheet link is versioned (style.css?v=2026-09-28): CSS is
  cached for an hour, and the new markup must not meet the old styles.
- README: nine pages, the shared structure, the CSS-version rule, and the
  operator's anti-slop waivers (the hero's volt rule, Space Grotesk).

Copy is unchanged apart from nav labels (the brief's legal constraints
hold). Still zero JavaScript and zero external requests. Impeccable
detector (controls 7/0, static + rendered at 1280 and 390, light and
dark): 200 unwaived -> 54, all of them the two waived brand devices;
two identical runs. Operator: "ship the site cleanup, keep the rule and
the font".
2026-09-28 16:29:02 -07:00
vh 2296fbba12 memory: Bonsai fork pin has a second copy on the smithy NAS 2026-09-28 15:44:44 -07:00
vh 60e6ed5b4e docs(fv-ml1): /tank/aimodels/mlx and the PROVENANCE-beside convention; Bonsai weights acquired 2026-09-28 15:43:52 -07:00
vh fd20183cbb feat(blender): extension set live in the GUI; MCP acceptance 9/9, probe made safe-mode compliant 2026-09-28 15:08:22 -07:00
vh 056433555f memory: Bonsai MMVQ-threshold follow-up (N=8 1.06x -> 1.21x); build image removed 2026-09-28 14:33:39 -07:00
vh 841f05de7a memory: Bonsai spike result on fv-ml1 GPU 3; blender-run --cpu 2026-09-28 14:23:21 -07:00
vh a51188c76a feat(blender-run): --cpu runs with no GPU attached (modelling/IO jobs stay off GPU 3) 2026-09-28 14:23:12 -07:00
vh 5e6a6f3b26 feat(blender-run): fixed container hostname fv-ml1-blender, so gethostname() is stable provenance 2026-09-28 14:00:07 -07:00
vh 59d436cc55 memory: first nightly restic run on repository-file verified all-green on all 8 hosts 2026-09-28 12:54:29 -07:00
vh 339de3cbcd docs(blender): SurfacePsycho STEP round trip into build123d verified by draupnir 2026-09-28 12:53:36 -07:00
vh 83dc497b40 feat(blender): pinned extension set in a read-only System repo, for the GUI and blender-run --extensions
Blender is now a mandatory stage in draupnir's pipeline (Prime, 2026-09-28), and draupnir asked
for eight add-ons from extensions.blender.org: SurfacePsycho 0.10.4, CAD Sketcher 0.32.1,
3D-Print Toolbox 1.4.1, STEP Importer 1.2.1, Bool Tool 2.1.0, LoopTools 4.7.7, MeasureIt 1.8.4,
3MF Import/Export 2.7.7.

- stacks/blender/extensions.lock pins each by version and archive sha256.
- scripts/blender-extensions sync builds fv-ml1:/tank/blender-extensions/5.2/system with Blender's
  own install-file, pre-warms and byte-compiles it, checks a read-only enable, then swaps it in.
  It refuses while the GUI or a blender-run job holds the old directory.
- conf/scripts/startup/fleet_extensions.py enables every package in the System repo: in a timer
  in the GUI (after the prefs load), and as --python ahead of the caller's args in
  blender-run --extensions (a failed enable exits 1 before the caller's script).
- It also patches SurfacePsycho's sp_overwrite_segment_selection from eval() to literal_eval():
  the eval walked past MCP safe mode (control: unpatched ran code, patched refuses).
- blender-run: --extensions (bind mounts via --mount so a missing source fails instead of being
  created); USER/LOGNAME set, which CAD Sketcher's getpass needs.
- compose.yaml mounts the repo read-only and the hook into the GUI container. NOT yet deployed.
- scripts/blender-probes/extensions_acceptance.py: one operator run per add-on, safe-mode
  compliant. Headless 8/9 online and with --network none; CAD Sketcher sketching is GUI-only.
  A Python audit hook saw no network/process events (positive control fired).
2026-09-28 12:50:57 -07:00
vh 7bd00ae77a feat(phasefinal-web): expand site to a multi-page cluster
Single page grows to nine: custom app development, systems engineering,
hosting, and AI services each get a detail page, plus engagement,
track-record, contact, and privacy. Cross-linked nav and footer on every
page; brand CSS extended in place (still zero external requests, zero JS).

Motivation: an app-store registration review flagged the single-page site
as minimal content. This version is intended to be reverted after approval.
2026-09-28 09:59:43 -07:00
vh 1756b891c6 memory: snapshot — Zigbee2MQTT + HA-MQTT route fix, Blender on GPU 3 (MCP per task + blender-run), semif 0.1.4, Worldtree reward config live and pushed; 3 entries archived; next: 2026-09-28 restic freshness check 2026-09-28 09:57:41 -07:00
vh 6f0c480693 memory: Blender access stays on the shared fleet login; batch path blender-run (Prime) 2026-09-28 08:50:09 -07:00
vh 46276611a7 fix(blender-run): unique remote job dir per local path, mirrored with --delete
draupnir review: the remote job dir was keyed on the basename alone, so two
local dirs with the same name shared one remote dir, and a rerun inherited
stale files. The dir is now <basename>-<8 hex of sha256(abs path)>, and it is
mirrored with rsync --delete, confined to that one directory. There is no rm on
a computed path. Verified: a file deleted locally is gone from the rerun's
remote dir. The /defaults stderr line is documented as harmless noise.
2026-09-28 08:43:55 -07:00
vh d0f68a3b18 feat(blender): blender-run one-shot headless wrapper + FLEETTOOLS entry
scripts/blender-run launches each call as a docker run --rm of the Blender
image on fv-ml1 GPU 3, capped at 64g / 48 CPUs. It needs no desktop and does
not affect the GUI container's lifecycle. --job DIR stages a local directory
to /tank/blender/jobs/<name>/, runs Blender with that as the cwd, and copies
results back. It always passes --python-exit-code 1, because Blender otherwise
exits 0 when a --python script raises (measured).

Tested headless: Cycles GPU and CPU, EEVEE via EGL, Workbench, an STL
round-trip, and exit codes (3, 7 and 1 pass through). There is no STEP
importer. Written for draupnir's design work, and indexed in FLEETTOOLS with a
detail file.
2026-09-28 08:39:43 -07:00
vh 2b38cfd1b0 memory: worldtree-instance-configs pushed (Prime) 2026-09-27 15:05:15 -07:00
vh 179309df3f memory: Worldtree reward config live on demo+personal; Blender MCP registration is per task (Prime) 2026-09-27 14:39:43 -07:00
vh ac1cd29afa feat(blender): agent control via mcp-for-blender (in-container, ssh stdio)
The MCP server (mcp-for-blender 2.1.1, frozen requirements) runs inside the Blender
container. Its add-on is vendored at upstream 41a18432 (MIT) and started by a
startup hook. scripts/blender-mcp carries the stdio over ssh + docker exec, so the
add-on socket, which runs arbitrary Python with no auth, stays on the container's
localhost with no published port. It also runs there because viewport screenshots
need a filesystem shared by server and Blender. Telemetry is off and safe mode is
on. The hook also defaults Cycles to OptiX on GPU 3, because safe mode forbids
agents from touching preferences.

Verified end to end from nh3-dev: 36 tools; a GPU render of an agent-built scene;
a viewport screenshot; and safe mode refusing 'import os'. Blender left down
(on demand).
2026-09-27 14:34:07 -07:00
vh 7ae7193c21 feat(blender): Blender 5.2.2 LTS on fv-ml1 GPU 3, on demand (Prime)
Uses linuxserver/blender (Selkies Wayland desktop, NVENC), digest-pinned. The
container sees only GPU 3, has restart "no", and runs only while in use,
because GPU 3 is the reserve card for a full-size vLLM seat. Web desktop on
:3001 with basic auth (vault fv-ml1/blender-web-password). /work is on /tank
and is not backed up; /config lives under /opt/docker (restic).

Acceptance: Cycles finds the card on OptiX and CUDA (sm_120 kernels ship in the
build). The heavy self-test renders in 3.54 s on OptiX vs 22.71 s on CPU, a
functional check with n=1. Web auth answers 401 without credentials and with a
wrong password, and 200 with the right one.
2026-09-27 13:57:35 -07:00
vh 0b8632a7ed fix(esh-docker-vm): /32 route to Home Assistant over macvlan-shim
The shim holds 10.0.50.47/24, which gives two equal connected 10.0.50.0/24 routes,
and ens18's wins. Host-to-HA traffic therefore left via the macvlan parent and was
dropped. HA lost MQTT to the broker on this host on 2026-08-19, 2026-09-21 and
2026-09-25 (the last lasted two days). This adds an ifupdown if-up.d hook that
routes 10.0.50.46/32 via macvlan-shim; /etc/network/interfaces is not edited.
Verified: the route resolves via the shim, the host pings HA, HA reaches :1883,
and HA reconnected to the broker. Diagnosis by ha-dev.

Also corrects the zigbee2mqtt acceptance note, which had wrongly reported HA as
connected.
2026-09-27 12:50:05 -07:00
vh c7b32418e1 feat(zigbee2mqtt): Zigbee2MQTT 2.14.1 on esh-docker-vm, replacing HA's ZHA
Requested by ha-dev; approved by Prime in this session. Radio: SLZB-MR1U chip 0
(EFR32MG21, EmberZNet 8.0.2) at tcp://10.0.90.10:6638, adapter ember. A fresh
network was formed on channel 25, PAN 0xCFF4. The broker is reached as mosquitto
user zigbee2mqtt, with HA discovery on homeassistant/. The frontend on :8099 is
token-protected.

State and the network key stay host-only in /opt/docker/data/zigbee2mqtt
(root 0700, restic). The repo carries only compose, .env.example and the README.
configuration.yaml refers to the secrets as !secret.yaml, and those references
survived Z2M's v4->v5 settings migration.
2026-09-27 12:43:49 -07:00
vh 47cad33dd1 fix(semif): 0.1.4 — object states ending in ) ; } no longer 422 (INV-7)
SemIf's shared scorer trims one token at the state boundary. When an object
state's last value ends in ')', ';' or '}', the JSON that follows re-merges two
tokens back, so score_shared refused the request with 422. The engine now wraps
semif_phase1.shared._state_prefix to keep only the tokens the full prompts
share. Each row scores the same token sequence; only the prefill/suffix split
moves.

Startup proves the fix is in effect, not just installed (heid bug hunt SKAL,
folded). It checks that the hook is callable and is what score_shared resolves,
that an ordinary state keeps upstream's whole prefix, and that a merge-prone
state scores through the shared path.

Real tokenizer: 154 states, 23 refused before and 0 after, with no ordinary or
authored144 prefix changed. Acceptance: 144/144 parity. Shared vs direct 71/72;
the miss is a bf16 tie that flipped across a plain restart (see README).
2026-09-27 10:23:14 -07:00
vh 7e11cf247b spike(semif): SemIf as Cicada's mood source is slower and less apt (no service change)
Against talk /face's guided pose (first paragraph 246 ms median), SemIf in
parallel adds 32 ms and SemIf-first adds 94 ms (n=72 each, noise floor 16.5 ms).
Removing the pose header saves only ~31 ms, and SemIf shares GPU 1 with the LLM.
Acceptable pose 67% vs 92% on clear-emotion lines, and the mood carried through
mundane follow-ups 7/15 vs 14/15. SemIf gestures far less (13% vs 58%).

README: rotations cost options^2 in suffix tokens, and /decide/shared returns
422 when an object state's last value ends in ) ; or }.
2026-09-27 09:47:28 -07:00
vh 89301a89fe docs(litellm): correct scalar-judge passthrough access semantics
The comment claimed 'gateway-key-gated'; v1.97.0 actually gates auth=true
passthrough routes on per-key metadata allowed_passthrough_routes (OSS path,
not Enterprise), 401 unauthenticated. Grants applied live via /key/update for
all-agents-local and the worldtree gateway key; comment-only change here,
picks up with the next conf deploy.
2026-09-27 09:29:50 -07:00
vh e268ff7c99 spike(semif): consumer fit for Wyrd scene change and Cicada affect gate (no service change)
Harness consumer_fit.py runs a consumer's per-turn decisions over hand-labelled
cases, with rotations, a content-free null control and tagged positive controls.

Cicada: an input-only 'does this earn a visible reaction?' gate scored 30/31
with descriptive options and 19/31 with terse yes/no options. Scoped by
Cicada's 2026-09-20 ruling (affect is emitted once, no mood-ring classifier).

Wyrd: on 3 real seed graphs, the first place-change wording failed its
positive controls (1/6 moves). A location-anchored rewording scored 21/21, and
exit selection scored 18/21. Semif fits the choice, not writing the node.
2026-09-27 09:20:35 -07:00
vh 3048e4194f memory: snapshot — semif 0.1.3 live (averaging + fast kernels + bug-hunt folds), hermes-gateway restart for SVOS seat_up, overnight backups green; next: SemIf spike (scope to confirm with Prime) 2026-09-27 09:00:51 -07:00
vh 77b8cb449c feat(semif): 0.1.3 — order averaging, fast kernels, bug-hunt hardening (Prime)
Order averaging (Prime, after the 739aa03 spike):
- A decision may set orderings: rotations|all (all only for <= 4 options). Every
  ordering goes to the engine in one shared batch.
- The reply keeps each native result and adds combined {probabilities (log-mean),
  top, agreement, spread}.
- Through the service on SemIf's labelled sets (252 rows): 78.6% -> 88.1%
  (group-bootstrap 95% CI +5.1..+14.3). Unanimous agreement is 94.5% accurate.

Fast kernels: flash-linear-attention 0.5.2 and causal-conv1d 1.7.0 are now the
default build. A/B on the empty GPU 3:
- parity with upstream went from 142/144 to 144/144;
- a ~2k-token /decide went from 169 to 92 ms server-side;
- short 3-rotation batches cost ~3-6 ms more.
triton builds a C shim at runtime, so the image carries gcc. Without it the
warm-up failed and startup failed closed.

Heid bug-hunt panel (4/4 arms, thread 01M3H3F4RR7XBP90KQ3A39H4SX), folded:
- Startup validation: VRAM cap 0 no longer means uncapped (C1); limits must be
  >= 1 (S1); the token must be visible ASCII (S2); the calibration file must
  exist and parse, with T in [0.05, 20] (S8, and C3's NaN leg).
- The body limit is checked before a chunk is kept, and a Unicode-digit
  Content-Length no longer crashes (C2, S3).
- Failures while building the response now get the 500 envelope (C3).
- 429 busy past SEMIF_MAX_QUEUE requests in progress (C6).
- The engine releases memory on every non-validation failure, unchained after
  gc; an empty OOM message is handled; 'out of memory' RuntimeErrors map to 503
  (C4, C5, S9).
- The entry point forces HF_HUB_OFFLINE (S10). README wording fixed (S5, S6).
- New guard tests close the gaps the arms' mutation grids exposed: early stop of
  the body read, a shared-route lock, calibration pass-through, the gc cycle,
  the exact caps, TorchEngine.load's arch and device checks, and the offline
  entry point.
86 tests.

Deployed on fv-ml1 GPU 1: parity 144/144, OOM and burst release verified, shared
capacity 63/51/26/16 rows at ~140/520/1960/3900 prefix tokens.
2026-09-27 03:27:15 -07:00
vh d7ad235365 memory: snapshot — semif live + averaging spike (build 0.1.3 + fast-kernel trial next), restic creds out of units on all 8 hosts, infra-ops on vm-esh-nas, augaman fv-ml1 instance removed; 4 foot-guns 2026-09-27 02:55:39 -07:00
vh 739aa03123 spike(semif): latency profile and order-averaging measurement (no service change)
Latency, measured from nh3-dev (3 runs x 20 per condition; network floor 31 ms):
- /decide short: 71 ms end to end, 38 ms server-side;
- /decide with a ~2,000-token state: 210 / 169 ms;
- shared, 3 rotations: 113 / 79 ms;
- shared, 6 orderings: 137 / 99 ms.
Qwen3.5's fast kernels (causal_conv1d, flash-linear-attention) are not installed,
so transformers falls back to its reference PyTorch paths. That is a speed lever,
and using it needs a parity re-check.

Averaging over option orderings, on SemIf authored144 + perturbations108 (252 rows,
72 groups):
- a single ordering scores 78.6%;
- log-mean over the 3 rotations scores 87.7% (+9.1 pts, group-bootstrap 95% CI +4.7
  to +13.8);
- all 6 permutations score 88.1%.
Rotations capture nearly all of the gain. Rows where the rotations agree unanimously
(161) are 94.4% accurate; split rows (91) are 75.8%.
2026-09-27 02:50:14 -07:00