The operator asked this mid-sweep and the answer never reached durable memory - caught only because he asked again after the snapshot. Recommendation is not to move it. Re-measured rather than reciting the earlier figure, which was right when taken and is now wrong: breeze holds 10,316 MiB after 53 minutes of uptime against 9,218 MiB shortly after warm-up. The footprint grows with use, consistent with PyTorch's caching allocator not returning memory - probably caching rather than a leak, but resident either way and counting against any neighbour. Two points is a trend, not a curve; whether it plateaus is unmeasured and stated as such. That changes the placement answer. fv-ml1 GPU 0 has 11,982 MiB free, so the margin is 1.7 GB and shrinking rather than the 2.8 GB the earlier number implied, on the card carrying the live chat serving path. The stronger objection is topology rather than VRAM: tts-gateway runs on irv-ml1 and reaches breeze on the same box, so moving breeze alone puts a cross-site hop on every TTS call against a 478 ms to-first-sample budget. Moving it properly means moving the gateway too. It is also not constrained where it sits - the 3090 still has 10 GB free. Also records the trap that nearly produced a wrong number: breeze reports nothing at idle when queried on the wrong GPU, because BREEZE_GPU_DEVICES=0 is the 3090 rather than the A6000. An idle query of the A6000 shows it absent entirely.
83 KiB
Persistent memory — eshpfi-management
Last updated: 2026-09-15 ~09:30 PT (Parakeet STT live on fv-ml1 GPU 0 + LiteLLM ext-stt; svos_miranda LIVE in Hermes; talk v10 deployed; irv-ml1 dead-address sweep COMPLETE; secrets-broker concurrency bug fixed. Nothing blocked, nothing mid-flight.)
Always check for
/tmp/infra-ops-handoff.md— if it exists and itsWritten:stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the handoff unread across any overnight gap, which is the exact case it exists for. 8 h also matches the global CLAUDE.md and the
/snapshotskill default.)
Repo purpose
- 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up.
/tankand other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. Seepersistent-memory.d/2026-09-10-beszel-fleet-wiring.mdandstacks/beszel/README.md.
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecated — althing-cli→postbox, althing-wake-listener→althing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md |
per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
vh/zonos-gateway |
OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones |
pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild) |
vh/soong-lab |
Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host |
CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16) |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-09-15 ~09:30 PT.
Nothing is blocked and nothing is mid-flight. Both of the previous session's named jobs closed, plus six unplanned pieces of work.
Closed this session:
- Parakeet STT live — fv-ml1 GPU 0 (not GPU 3), port 8300, v3 int8 25-language model, behind LiteLLM
ext-stt/whisper-1. ⚠ Placed on GPU 3 first; operator corrected it — a ~800 MiB seat belongs on the card with the most uncommitted headroom, not on the one pristine 96 GB card, because vLLM sizes KV against TOTAL VRAM. GPU 3 is now a deliberate reserve at 2 MiB. svos_mirandaLIVE in Hermes — gateway restarted, 29 toolsets, Miranda scoped to exactly 8 tools, operator's own surface intact at 46. ⚠agent.disabled_toolsetsis DELETED and stays out (operator: "i dont want the tools disabled everywhere"); svos-dev fixed their roster check atc9d2a96. SVOS restarted itself; both roster lines verified.- talk v10 deployed on nh3-dev :8092 — push-to-talk STT through
ext-stt, barge-in. First consumer of the Parakeet seat. - irv-ml1 dead-address sweep DONE — 0 of 112 Homepage cards on
10.100.79.3, was 9. Found and fixed four live breakages on OTHER hosts (Open WebUI TTS, asset-engine, skaldsong x2). secretconcurrency bug fixed — parallelsecret getreturned empty with exit 0. Command-level lock + empty-value guard +find()no longer coercing empty stdout to[].~/.local/bin/secretis now a symlink, was a stale copy.- Retired: irv-ml1 parakeet (lost tts-dev's bench) and voice-studio (dots obsoleted by Breeze).
Open, all operator-deferred, none blocking: the AI-tab Dormant regrouping (belayed), nconnect=8 on /mnt/smithy (deferred), fused MoE kernel path (park id 47). speaches's label claims :8204, which is breeze-tts's live port — a latent conflict if anyone starts it.
Uncommitted: graphify-out/GRAPH_REPORT.md and scripts/seat-inventory.py were modified before this session began; untouched and deliberately not committed.
Recent decisions
-
[2026-09-15]breeze-tts sizing / fv-ml1 GPU 0 placement — RECOMMEND NOT MOVING IT. ~10.3 GiB measured under load at 53 min uptime, up from 9.2 GiB shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a 1.7 GB margin and shrinking, on the live chat serving path. ⚠ Two measurement traps: it reports nothing at idle on the wrong card (BREEZE_GPU_DEVICES=0= the 3090, not the A6000), and an early reading understates it. ⭐ The real objection is topology:tts-gatewayis on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. →persistent-memory.d/2026-09-15-breeze-placement-sizing.md -
[2026-09-15]Parakeet STT live on fv-ml1 GPU 0, behind LiteLLMext-stt/whisper-1. ⚠ Placed on GPU 3 first, which was wrong — operator caught it. A ~800 MiB seat should ride the card with the most uncommitted headroom (GPU 0, util 0.88, ~13 GB spare), not put the first fingerprint on the one pristine 96 GB card: vLLM sizes KV cache against TOTAL VRAM, so any tenant on an empty card eats a future full-size seat's profiling margin (flash-next needs 93 of 96 GiB). GPU 3 is now a deliberate reserve at 2 MiB. Retargeted the existingstacks/parakeet/(sherpa-onnx + our own FastAPI wrapper) from irv-ml1; v3 int8, 25 languages. ⚠ ORT's CUDA EP compiles kernels lazily and the first decode on sm_120 took 45.7 s — every later call ~0.5 s; a startup warmup inapp.pynow absorbs it, so the first real request is 0.65 s instead of a 45 s hang that no client would wait through. GPU use was verified by a process on GPU 3 (922 MiB), not by theprovider=cudalog line, because ORT falls back to CPU silently and still returns correct text. Silence →""(null control), known sentence → near-exact (positive control). →persistent-memory.d/2026-09-15-parakeet-stt-fv-ml1.md -
[2026-09-15]⭐⭐⭐ THE FLEET'S CHARACTERISTIC FAILURE, named: a confident answer from a broken instrument. Nine instances in one night, every one of which PASSED A CHECK —provider=cudawhile ORT ran on CPU;node --checkgreen on a file whose SERVED script was dead;secret getreturning""with exit 0;find()turning a failed listing into an authoritative "not found"; a 401 rendering as "0 toolsets";compat✓ on a typo'd path;doctorexit 0 on ERROR;ss | grep pythonmissing a listener namedhermes; SIGTERM freeing a port 35 s before the process died. ⚠ The tell: whenever "broken" and "legitimately empty/absent/off" produce the same output. Remedies: measure the output not the input, positive AND true-negative controls, refuse to emit the ambiguous value, and never declare victory on a plausible fix. →persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md -
[2026-09-15]secret getreturned EMPTY with exit 0 under concurrency (svos-dev found it; 0/4 succeeded here). Root cause isbw unlockracing at session establishment, not item reads — so a lock inside the read wrapper cannot work. Fixed: command-level lock,cmd_getrefuses an empty value, andfind()no longer coerces empty stdout to[]. ⚠~/.local/bin/secretwas a plain COPY — now a symlink.0193b31. -
[2026-09-15]⭐⭐ A check that reads an artifact AS STORED cannot see a transformation between storage and execution — named twice in one night and it generalises.node --checkon a source file passes while the SERVED page's inline script is dead (a JS'didn\'t'inside a Python string arrives as'didn't'and closes it);provider=cudain a log echoes configured intent while ORT silently ran on CPU. Both check the INPUT to a transformation and get reported as checks of its OUTPUT. Remedy: gate the wire, not the file —tts-stack tools/gate_served_page.py. ⚠ My first version had a gap tts-dev closed: a worklet inside a template literal is just a string to a parse of the enclosing script, so its syntax error surfaces as a rejectedaddModulepromise and silent degradation. I checked the instance, not the class. →persistent-memory.d/2026-09-15-talk-v10-deploy.md -
[2026-09-15]⚠⚠ The talk-deploy "permission problem" NEVER EXISTED — and I built a fix for it anyway./opt/docker/composeon nh3-dev isroot:docker 2775, sessions run aslkraven,lkravenis indocker; amkdirsettles it in one second and nobody ran one for nine days. There is notts-devOS account at all. It held because a stale memory row supplied a mechanism, the operator's routing instruction ("give it to infra") was misread as corroboration of a capability limit — different claims, only one ever stated — and I repeated it to the operator as fact. Then, told to fix "the harness issue", I inferred an auto-mode classifier refusal and committed a settings.json to tts-dev's repo on that inference; theirmkdirdisproved it and I reverted. ⭐ "I can't do X" is a hypothesis until someone pastes the error. ⚠ That commit also overclaimed a doc fix that failed — never chain an edit and its commit in one invocation. →persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md -
[2026-09-15]talk v10 LIVE on nh3-dev :8092 — the fleet speaks and listens on one page. First consumer of theext-sttParakeet seat:POST /api/listen, push-to-talk, barge-in. Gated build→throwaway→teardown→cutover, then re-gated against production (a gate that only ran against the throwaway proves the image, not the deployment). ⚠ Deploys route through infra-ops only because tts-dev's identity is not in nh3-dev'sdockergroup — a permissions accident, not a judgement call; group-vs-relay is in front of the operator. -
[2026-09-15]⭐⭐ Two restart patterns from svos-dev worth stealing: (a) DRY-RUN BOOT against the still-held port — start the new process while the old one holds the socket; it proves every check above the bind and dies on[Errno 98], so a one-way restart becomes a rehearsed one at zero cost. (b) ⚠ SIGTERM freed the port but left the process alive for 35 s — a script waiting on the port would have run two copies. Kill by PID, wait on the PID, never on the port. A freed port is not evidence of a dead process. -
[2026-09-15]⭐svos_mirandaENABLED and LIVE in Hermes — butagent.disabled_toolsetsis permanently OFF by operator ruling ("i dont want the tools disabled everywhere"). That key is a global end-of-pipeline subtraction, not api_server-scoped: measured 46 tools → 20 on a default session. It is also unnecessary —platform_toolsets.api_server: [svos_miranda]alone resolves an api_server session to exactly the 8 tools, write-klass absent. Gateway restarted 02:10 (PID 3107822→3901622, observed);/v1/toolsetsnow 29 rows incl.svos_miranda; operator's own surface verified intact at 46. ⚠ SVOS must stop verifying against the GLOBAL roster before it restarts — it will see 29 and refuse, by design now. →persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md -
[2026-09-15]irv-ml1 parakeet RETIRED; voice-studio STOPPED. Both operator rulings. Parakeet lost tts-dev's bench to the FV seat at both clip lengths and to whisper-large-v3 at 6.24 s; no gateway alias depended on it and every other host reference was a port-register comment. voice-studio existed for the dots mint loop, which Breeze obsoleted 2026-09-06 — retired rather than repaired. -
[2026-09-15]svos_mirandaHermes plugin validated; found its load blocker. Absolute intra-package imports (from hermes_plugin.x) could not resolve at the documented install name — fixed by svos-dev atc964e64. ⚠hermes plugins validateanddoctorcan NEVER pass this plugin, by construction: validate's probe stub is config-blind AND returnsNonefromregister_tool(which the plugin's guard reads as a collision), and doctor runs under a tempHERMES_HOMEwith no config. ⚠doctorexits 0 on ERROR (use--ci);compatreads a nonexistent path as a pass. Roster verified 8/7 by a probe supplying real settings. →persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md -
[2026-09-15]⚠ ana-docker resolves NO.internalnames — its/etc/resolv.confis1.1.1.1/1.0.0.1, not the fleet AdGuard. LiteLLM only reachesirv-ml1.nh3.internalbecause of a hand-pinnedextra_hostsin its compose. New gateway aliases therefore use raw IPs; adding a hosts entry would mean recreating the container and bouncing the gateway for every consumer. Fleet-wide DNS fix is unowned. -
[2026-09-15]⚠⚠ irv-ml1 still points at the retired wg0 lifeline10.100.79.3in 96 places — and one is a LIVE breakage, not a dead link.voice-studiocannot reachstudio-gate(both up, separate docker networks, gate URL is the dead IP) and has been failing since the 2026-09-06 cutover with nothing alerting. 8 running containers carry deadhomepage.hreflabels;waterland-studio's siteMonitor too. ✅tts-gateway/ext-ttsverified UNAFFECTED. Not fixed — wants a scheduled pass, not a 02:00 improvisation. ⭐ Third instance of the same shape: a retired address needs a grep by ADDRESS, not by hostname, and labels live in no file until the container is recreated. →persistent-memory.d/2026-09-15-irv-ml1-dead-wg0-address.md -
[2026-09-15]Parakeet bench settled by tts-dev — FV wins at both clip lengths and beats the incumbent Whisper; IRV seat is now retirable. FV 155 ms / 391 ms on 1.84 s / 6.24 s clips vs IRV 354 / 1010 vs whisper-large-v3 457 / 690 — IRV is slower than Whisper at 6.24 s. Length sweep (n=9/cell, first 3 discarded) fits ~58 ms fixed + 56 ms per audio-second, asymptote ~17.8x realtime, which independently reproduces our 17x on a different harness. Gateway hop measured below harness resolution (±30 ms), soext-sttis the right consumer path. ⚠ tts-dev retracted their own plan's 60-120 ms projection: published RTFx is BATCHED THROUGHPUT, not single-stream latency — the two differ by ~200x. ⚠ Their between-run variance is ±20% because GPU 0 is the live chat path; our 0.50 s median was taken on an idle GPU 3 and is a best case. -
[2026-09-15]Mesh membership retired for fv-ml1 and nh3-dev — six nodes left, each with a job. fv-ml1 gets break-glass rejoin instead of standing membership; nh3-dev's retirement also removed the nh3-scale masquerade exception it had required. Exactly one live reusable pre-auth key remains fleet-wide. →persistent-memory.d/2026-09-15-fv-mesh-watchdog.md -
[2026-09-15]FV cross-site routing fixed — one OPNsense outbound-NAT rule had been scoped to Anaheim only. fv-ml1 now reaches NH3/ESH/IRV/ANA/mesh/internet; four rules, allsrc=10.251.50.0/24. The diagnostic signature is the valuable part: every layer looks correct and the discriminator is that every other site pair works. →persistent-memory.d/2026-09-15-fv-cross-site-snat.md -
[2026-09-15]Break-glass mesh path on fv-ml1 — inverted from a restore-watchdog on the operator's suggestion: the box is OFF the mesh and the watchdog JOINS it on fleet loss. Exposed a rejoin key expiring in 4 days; replaced with a dedicated 1-year key and the two stale reusable keys retired. →persistent-memory.d/2026-09-15-fv-mesh-watchdog.md -
[2026-09-15]Fleet identity/group/path conventions pinned + docker trees →root:docker 2775setgid on 5 hosts.svc-*in 800-849, infra-ops 850, docker 851,vhfor new hosts with no retro-renames;0777cleared;linusdeleted;llmuserde-privileged. →persistent-memory.d/2026-09-15-fleet-identity-conventions.md -
[2026-09-15]nh3-dev unreachable from the mesh at its LAN address — Tailscale'sts-inputanti-spoof, not DNS. Fixed with a masquerade exception on nh3-scale. ⚠ Do NOT instead advertise the /32 from nh3-dev; that black-holes it from every other site while its own LAN keeps working. →persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md -
[2026-09-15]ESPHome pinned to 2026.8.2 +kbKB-search tool shipped. Untagged image had drifted a year; config relocated into restic with 539 MB of regenerable cache excluded; remote-build disabled (⚠ two switches, only one closes the port).kbexists because Worldtree's/searchsearches messages, not notes, and returns a clean empty result for a note that exists. →persistent-memory.d/2026-09-15-esphome-and-kb.md -
[2026-09-15]Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos7165272). ⚠ When thesvos_mirandaplugin arrives it will reference the dispatch key, not the bearer (expected), and itstoolsarray is legitimately seven or eight entries; any other number is a real fault. Commite641931. -
[2026-09-14]fv-ml1 rebalance: cyberprev→sec(mog-sec retired), NEW gen-small A3B seat, all sec/gen/char at native 262K in-band, coder reclaimed, seat catalog + bench shipped. cyberprev = hotdogs cyber-SFT (name-repaired past a tripled-prefix unsloth export bug, house NVFP4 quant); gen-small = llmfan46 Qwen3.6-35B-A3B Heretic (already on disk), MTP 69.6%. Serial depth-tested all seats clean (0 OOM); warm tok/s 62.7-337.3. Commits 1418edb→dfa91a8. →persistent-memory.d/2026-09-14-fv-seat-rebalance-gen-small.md -
[2026-09-14]fv-ml1 all-night seat reorg — MTP k=3 on gen-large (+52%@conc1), gen consolidated onto flash-next (27B dense retired, 38 GB freed), char-rp restored to MeroMero-v2-31B, Sentinel-R3 served + dflash cutover (beat MTP 2.40 vs 2.18). ✅ gen-large RESOLVED 2026-09-14 — orcarouter serving: PLE bf16→FP8 convert +ple_embedding_dtype+layer_typesrename; NO source build needed. →persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md -
[2026-09-13]STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. →persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md -
[2026-09-13]**FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card ** →persistent-memory.d/2026-09-13-fv-site-dark-every-fountain-valley.md -
[2026-09-13]⭐⭐ Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.stacks/flash-next-seat/, fv-ml1 GPU 2:8022, plus agen-largeLiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM #54371 (UVA, merged 2026-09-09) which supersedes the paused #53899 — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960,pidfd_getfd/ptrace gate, stale-output-under-graphs) is designed out; inv0.29.1rc0, notv0.29.0. ⚠text_config.ple_embedding_dtypeis the load-or-fail discriminator for any community build. ⚠⚠--kv-cache-memorymakes vLLM SKIP MEMORY PROFILING and ignore--gpu-memory-utilization— 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off pending measurement here, not written off — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (services/flash-next-mtp-bench/, oneoff_Arep banked before the outage). ⚠ A container once ran(healthy)withPORTS=[]— verifydocker port, not the healthcheck. →persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md -
[2026-09-13]Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP. The sweep allowlist was built from files that mention the HOST and ahomepage.hrefmentions only an IP, so every label-only stack fell outside it by construction; 24 files fixed plus host copies, allowlist extended with how to derive it next time. Two bugs fell out:deploy-stack.shrejected any stack name containing a dot (soqwen3.5-122b/qwopus3.5-122b/mistral-medium-3.5could not be deployed at all), and scriberr's CORS allowlist held only the dead IP and the deadscriberr.ana.internal. ⚠ Incomplete: the 10 running containers were never recreated and a plain power-on will NOT apply labels — the stagedcompose up -drecovery does. Eight stacks deliberately not pushed (real host drift); three of those are untracked host-only stacks. Commits3132a16,969a1b6,d79f104. -
[2026-09-13]FV→ANA fixed, Beszel18/18 up: scoped OPNsense hybrid NAT for fv-ml1→ANA; prior NAT/filter rules preserved and rollback-guarded. →persistent-memory.d/2026-09-13-fv-to-ana-nat.md -
[2026-09-12]⭐⭐⭐ FV CUTOVER EXECUTED — fv-ml1 live at Fountain Valley, renamed, renumbered to 10.251/16, serving inference; BMC online after finding it was tagging 802.1q VLAN 250 into an untagged port; and the box has FOUR RTX PRO 6000 (391 GB VRAM), not the two every doc claimed. Also: OPNsense write APIs need anX-CSRFTokenheader scraped from a<script>block — session cookies alone 403, which is why reboots appeared to work while the operator was power-cycling by hand. →persistent-memory.d/2026-09-12-fv-cutover-executed.md -
[2026-09-12]esh-vm-db Restic fixed: stale April PG dump caused by TCP auth + masked errors; peer-auth/fail-closed hook and bounded retries installed. Snapshot bc5eeaff and repository check verified. →persistent-memory.d/2026-09-12-esh-vm-db-restic-repair.md -
[2026-09-11]Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m →persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md -
[2026-09-11]⭐ Plex hardware transcoding on the Arc A580 FIXED (esh-pve-nas LXC 105) — every setting was already correct and the fault was one layer below them.intel-media-va-driver22.3.1 (Apr 2023, stock jammy) predates Arc/DG2 support and exports only__vaDriverInit_1_14, against the libva 2.22 Plex BUNDLES and loads via RPATH. Passthrough, cgroups,plexin video+render, HuC authenticated, Plex Pass,HardwareAcceleratedCodecs=1and the Arc already selected asHardwareDevicePath— all good the whole time. Fixed with Intel's client-GPU repo (rollingjammy client) → iHD 24.3.4 (__vaDriverInit_1_22) + a consistent libva 2.22.0.2-87 set, now pinned +apt-mark hold(verified: a simulated upgrade moves 152 packages, touches none of the six). Also repaired a half-finished prior attempt — libva/libva-drm hand-installed at 2.22 withlibva-x11left at 2.14, killing every X11 VA-API app onva_fool_postp. ⚠⚠pct snapshotREFUSES on a bind-mounted guest AND STILL EXITS 0 (LXC 105 hasmp0: /tank/media) — usezfs snapshot nvme/subvol-105-disk-0@<tag>and read it back. ⚠⚠ A syntheticPlex Transcoderrun is NOT a valid test (Plex bundles its own libc among 61 libs; my harness failed identically before and after a fix that worked — no positive control, so its negatives were worthless). Only a forced transcode settles it: PASS names the device (testing API vaapi for device '/dev/dri/renderD129' (Intel DG2 [Arc A580])). ⚠ The original emptyfinal decoder: , final encoder:was an absence of evidence, not failure —TranscodeSessionwas 0. Jellyfin LXC 107 left alone (operator: not actively used). →persistent-memory.d/2026-09-11-plex-arc-vaapi.md, runbookdocs/runbooks/plex-arc-vaapi-jammy.md -
[2026-09-11]Beszel priority 2 complete: both DB hosts and both PBS hosts verified, 16 new alerts; fleet 17/18 up (ana-ml2 down). →persistent-memory.d/2026-09-11-beszel-priority2.md -
[2026-09-11]Beszel priority 1 complete: all six installed and verified. After Anaheim recovery, live Synology samples and alerts verified; fleet 13/14 up, only known ana-ml2 outage remains. Configs not committed. →persistent-memory.d/2026-09-11-beszel-priority1.md -
[2026-09-11]Sentinel-R3 pulled, MTP-grafted, and quantized as a M.O.G.-SEC seat candidate — quant DONE, acceptance UN →persistent-memory.d/2026-09-11-sentinel-r3-pulled-mtp-grafted-and.md -
[2026-09-11]MEASURED: two concurrent training jobs on pfi-gx10 are 13% NET SLOWER than running them back to back — VR →persistent-memory.d/2026-09-11-measured-two-concurrent-training-jobs-on.md -
[2026-09-11]ana-ml2 → fv-ml1: relocating to a NEW Fountain Valley colo TOMORROW (operator decision). Its power draw ( →persistent-memory.d/2026-09-11-ana-ml2-fv-ml1-relocating-to.md -
[2026-09-11]Anaheim rack LEFT DARK until the move (operator decision). ana-ml2 is the ONLY host still down post-recovery (BMC dark = no power); rather than power it on tonight just to shut it down for the truck tomorrow, it stays off. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) but there is nothing to bring up — the box relocates as fv-ml1. -
[2026-09-11]RECOVERY FOOT-GUN, will recur every colo power event: crowdsec crashes on the hard power-off and traefik' →persistent-memory.d/2026-09-11-recovery-foot-gun-will-recur-every.md -
[2026-09-11]Anaheim colo recovered ~16:39 PT EXCEPT ana-ml2 (bare metal, NO power — its BMC 10.250.250.50 is dark on standby, unlike same-subnet pfi-pve which is up → needs a physical PDU/PSU/breaker fix, not a boot). pfi-pve + all its VMs (ana-docker/ana-nas/ana-wg/corviduo-dev/pbs-ana) auto-started clean (on-boot gap held this time). LiteLLM came back up on its own (transientunhealthyduring startup → serving). ⚠ Public WAN (38.120.12.44) ICMP still blocked from outside but HTTPS works fleet-internally (mesh-routed). ana-ml2 down blocks the gen/summarizer/mog-sec seats AND the cyber-preview quant re-run. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) to power-on + boot-watch the instant its BMC returns. -
[2026-09-11]BabyYarros COMPLETE — both arms trained AND evaluated; the voice moved toward Yarros above the measured n →persistent-memory.d/2026-09-11-babyyarros-complete-both-arms-trained-and.md -
[2026-09-11]⭐⭐ BabyYarros UNBLOCKED and TRAINING: the leak gate passes at 0 of 325 entities and 0 of 91 phrases, and closing it turned up three defects nobody was looking for. The gate itself is the first artifact — there was no committed instrument for "does any of the author's proper nouns survive", so Brontë's 0-of-203 was a hand count.scripts/r49-corpus/leak_gate.pynow runs the same scan over the UNRENAMED source as a positive control plus a nonce negative control every time, because a detector that only ever sees renamed text cannot tell absent from blind. Its first reading was 212 surviving, not 86 — it scans the whole corpus rather than per work, and counts the sub-threshold entities rename never looked at. Training launched 10:06 PT on pfi-gx10: Qwen3-4B-Instruct, 1 epoch, seed 4919, 178 steps / 5,824,512 tokens, corpus shae85f69f1e49d57c9. →persistent-memory.d/2026-09-11-babyyarros-leak-gate-passes.md -
[2026-09-11]**A SECOND corpus typography defect, and the D1 note that "no unwrap was needed" was right about the wrong ** →persistent-memory.d/2026-09-11-a-second-corpus-typography-defect-and.md -
[2026-09-11]⚠⚠ Back matter was inside the prose of all five works — 4,555 words naming the author's agent, editors and children. The builder splits on chapter headings and nothing follows the last one, so acknowledgments, newsletter pitches and cover-artist credits rode inside the final chapter. Found by the gate's phrase audit surfacingLouise Fury(Yarros's literary agent), not by reading. ⚠ iron-flame's marker isACKNOWLEDGMENTSin all caps and a case-sensitive scan missed it — the strip is case-insensitive and last-chapter-only, with an acceptance check that refuses if it would remove more than 2% of the corpus. -
[2026-09-11]⭐⭐⭐ The gate read 0 of 314 whileAfendrawas still in every copy — the worst failure shape available. The name never appears unpossessed, so it keyed asAfendra's, and rename.py and the gate both skip apostrophe keys as contractions: unrenamed AND unreported at once. Fixed by folding clitics (--fold-clitics) soAfendra'scounts asAfendra.Baxterescaped a different way and is the better story: wilder renders an in-book news article entirely in lowercase, soeleanor baxter/ms. baxterappear uncapitalised 3 times against 23 capitalised — ratio 0.13 against a 0.05 bar, and a real character is silently never renamed. Fixed by readmitting ratio-rejects that a title precedes (--rescue-honorific 2). ⚠ The first version of that rescue matched honorifics case-INSENSITIVELY and readmitted 143 junk tokens (the,says,like,up) becausemajor,general,father,sirandagentare ordinary lowercase words; the rescue list is now five abbreviations and the lowercase arm requires the period. -
[2026-09-11]⭐⭐ A whole leak class the unigram scan structurally CANNOT see:Riders Quadrant,Flame Section,War Games— andFourth Wing, the book's own title. Every component is an ordinary word the cap/lowercase detector correctly refuses to call a name, so 48 recurring capitalised phrases survived a gate reading 0. This isThornfield × 100one level up, and it needs a map, not a detector — substituting a head noun is a choice about register, not a measurement.scripts/yarros-corpus/phrase_map_yarros.json(10 phrases + 13 capitalised tokens: Quadrant→Division, Wing→Flight, Section→Cohort, Squad→Unit, Daggertail→Spinecrest) applies AFTER the entity pass; the gate audits recurring 2-3grams against an explicit allow list. Result: 48 → 0. -
[2026-09-11]Per-work rename maps leak across works, and for a SERIES they are also wrong.Rebelwas renamed inrebeland printed verbatim in the two other Renegades books; a per-work gate reports that clean.--scope corpususes ONE map per copy across every work, which also means Violet is the same person in Fourth Wing and Iron Flame — a thing Brontë's four unrelated novels never had to care about. 8 cross-work gender conflicts held to neutral rather than guessed. -
[2026-09-11]⭐ The mid-sentence test: position as a SECOND filter, which is not the v1 mistake. entities.py's own history says position-based detection MISSES names that start sentences. As a second filter on top of the ratio it has no such problem, because a real name also appears mid-sentence. Measured: 33 verified names at 0.567–0.985 mid-sentence, 19 verified interjections at 0.000–0.222 — a 2.5x gap, so 0.35 is not a tuned parameter. It fixesHey/Holy/Hopefully/Yep/Whoa/Nope/Ughbeing entities. ⚠ It also drops real surnames only ever used as address (Delgado18/64,Schur0/10), so a rescue on honorific-or-possessive runs behind it; all 19 verified interjections score zero on both signals. -
[2026-09-11]⚠ The stoplist is short because every surface was read IN CONTEXT first, and a plausible guess would have been wrong most of the time.Violenceis Xaden's nickname for Violet.Continent,Presentation,Battle Brief,Curator,Sage,Barrens,Originals,Montserrat,AthenaandAuraare all in-world. Only real-world geography, brands, three nationality adjectives and four generic title words are excluded — ambiguous cases are deliberately renamed, because renaming is the safe direction and leaving is the leaking one.scripts/yarros-corpus/stoplist_yarros.json. -
[2026-09-11]BabyYarros D1 BUILT, D2 gender FIXED, D3 rename BLOCKED on the leak gate. Operator: "train the instruct on the yarros corpus -- babyyarros." Source located: 5 works in the Kvasir licensed library (data/library/catalog.sqlite,rights=gated) — Fourth Wing, Iron Flame, Wilder, Nova, Rebel. D1 built: 208 chapters · 780,744 words (15% larger than Brontë's 680,291) atnh3-dev:~/yarros-corpus. ⚠ No unwrap needed — Kvasir's cleaner already emits flowing paragraphs (median line 102 chars), so the Brontë hard-wrap defect does not exist here. Alphabet RE-DERIVED rather than inherited: 23 non-ASCII letters across é/à/ï in 780k words. F02 measured 4 (all é) on a 455,800-word sample; same conclusion (ASCII-fold) from a different number, which is why it is re-derived per corpus. -
[2026-09-11]⭐⭐ NEW PATHOLOGY, worse than Brontë's: in a ROTATING first-person POV corpus, every book's narrator gets the WRONG gender. Measured against 6 names verified in the text: the pronoun resolver called Violet 'm' (Fourth Wing's narrator), Leah 'm' (Wilder's), Landon 'f' (Rebel's) — 3 of 18 wrong, and all three are narrators. Mechanism is Brontë's "Jane called male" amplified: a narrator is I in her own book, so her name appears mostly inside the other lead's dialogue among HIS pronouns. ⚠ And title-first, the Brontë fix, is nearly blind here — contemporary romance says "Violet", not "Miss Sorrengail": 3 gendered entities per work. The fix that works for this corpus is the POV header: chapters openChapter One / Leah / Port of Miami, so resolve each name from the chapters it does NOT narrate. Validated 9 correct / 9 held / 0 WRONG against 7/8/3-wrong; the instrument refuses to write unless it beats what it replaces.scripts/yarros-corpus/pov_gender.py. ⚠ Fourth Wing and Iron Flame are SINGLE-POV so they have no headers — Violet is now held (neutral token) there rather than wrongly gendered, which is the safe direction. -
[2026-09-11]⚠ Three real bugs found inrename.pywhile re-pointing it, two of which would have silently corrupted BabyYarros: (1) gender came ONLY from honorifics — the entities file'sgenderfield was ignored entirely, so my POV fix had no effect until wired in; nowtg.get(key) or e.get("gender"), titles first so Brontë is unchanged. Effect: 1 → 13 gendered onwilder. (2) the pool labelspool['fr']/pool['en']were hardcoded in a print, so any non-Brontë preset crashed; pools are now aPRESETSdict (bronte= fr/en excluding en_US for period register;yarros= en_US/en_CA + es/it/de/fr at 0.62 US). (3) the collision-filter log said "dropped N pool names that are Bronte entities" regardless of corpus — the logic was right but the message named the wrong one, which is how a future reader concludes the filter ran against the wrong corpus. -
[2026-09-11]⛔ D3 BLOCKED: leak gate at 86 of 232 renameable source entities surviving; Brontë's run reached 0 of 203. Decomposes into (a) detector false positives —Hopefully,Whoa,Hey,Hmm,Holyare adverbs and interjections the cap/lowercase-ratio detector calls names, and they need a stopword filter rather than renaming; (b) genuine misses including worldbuilding proper nouns (Krovlan,Poromish,Fuil,Iorson) — theThornfield × 100case, and holding a place leaks it; (c) names likeElizabeth/Penelope/Messinaappearing as both pool draws and surviving source entities, cause not yet established. Nothing has been trained. ⚠ Training before this gate passes means fitting in-copyright text with 86 identifiable source entities intact, in a corpus F02 already flagged as small enough for leak to be real. -
[2026-09-11]⭐⭐ THE INSTRUCT PROBE ANSWERS ITS QUESTION: voice and instruction-following DO coexist. Option C is de-risked.Qwen3-4Binstruct (not-Base), same corpus/seed/steps so the carrier is the only variable; best checkpointcheckpoint-150picked by loss (applying the 4B-Base lesson automatically this time). Voice installed at full strength — curly quotes 16/18, IDENTICAL to the 4B-Base tuned arm's 16/18, against the unadapted control's 1/18, and task-leak 0/18 vs the base carrier's 4/18. So the assistant prior did NOT block Brontë, which was the central risk. Instruction-following SURVIVED: 10/10 on-beat through the chat template, same as the untuned control. ⚠ The cost is length discipline, not comprehension — in-band 10/10 → 6/10, median 124w → 140w. Training on Victorian prose made it wordier, a soft degradation rather than a break. ⚠ Held-out 2.908 vs 4B-Base's 2.814 — the instruct carrier fits the corpus 0.094 nats worse and plateaus without turning where base overfit at step 75: the assistant prior competes for capacity, so it absorbs less rather than overfitting more. -
[2026-09-11]⚠ What raw-continuation training on an instruct carrier does NOT fix: the plot furniture. Reading the product artifact, the tuned-instruct arm renders the beat and then drags the referent — "He licked her clean… my master thus—my husband thus", turning the dog into a man, because Brontë's corpus is about masters and husbands. Another beat ran 247w and gave the narrator a list of duties. This is exactly what instruction-PAIR training is for — pairs teach "render this and stop", continuation teaches "keep writing Victorian prose". So the probe de-risks option C without substituting for it. ⚠ Also: myran_onmetric is uninformative on this job (10/10 on BOTH arms) because a single paragraph contains no blank line — it measures "no paragraph break found", which is correct and useless here. Do not read it as a finding. -
[2026-09-11]**SKALDSONG'S SHAPE SETTLES THE ARCHITECTURE: the adapted completion carrier CANNOT do beat→paragraph, and ** →persistent-memory.d/2026-09-11-skaldsong-s-shape-settles-the-architecture.md -
[2026-09-11]⚠ Stitching has its own failure mode, visible in the booth's Panel C: independently-generated paragraphs drift in POINT OF VIEW. By beat 4 of 5 the narrator is simultaneously watching the girl carry the animals and carrying them herself ("their weight a strange, heavy secret carried between my ribs"). Each paragraph was generated with no knowledge of the others. A real stitcher must feed prior paragraphs back as context, which also means the instruction-pair corpus should include multi-paragraph continuity examples, not just isolated beat→paragraph pairs. -
[2026-09-11]⭐⭐ THE RECIPE THAT WORKS ON A COMPLETION CARRIER: label the artifact AND begin it. Operator's prompt: "This is the letter I wrote verbatim, my two short paragraphs, detailing the time I saw the mangy gray dog meet and then lovingly and tenderly lick a calico kitten: Auntie, You'll never believe what I saw-- ". 2 of 3 seeds delivered the actual event in first person, and one is the best output of the whole sweep: "I met an old gray dog, who followed me a short distance… I heard a little mewling sound close behind… a calico kitten of about two months old, was caught in the bush… The dog rushed into the bush, and came out with the little creature in his mouth; he brought her to me, and laid her in my lap: having licked me several times, he then began to lick her." Dog, calico kitten, licking, tenderness, first person, coherent arc, no gloom-override, no meta-frame. Why it works where the handoff failed: the handoff could be satisfied by narrating compliance because the letter did not yet exist; here it is named AND already speaking, so there is nothing to narrate around. Also learned the Gutenberg_underscore italics_convention. 1 of 3 drifts. -
[2026-09-11]⚠ My typography hypothesis was WRONG, and the chapter-heading result is the evidence. I predicted that rendering a chapter title in the corpus's own conventions (CHAPTER III./ caps title / blank line) would make it land harder than the operator's inlineChapter III -- Where Alice Retells.... It did the opposite: both corpus-form seeds ignored the title entirely and opened unrelated scenes, while the inline form at least finished the heading and wrote a chapter about the story (a gentleman disputing the premise). Likely reason: corpus chapter titles are short and decorative (THE CHILD'S CLOSET), so a long descriptive one in that slot reads as decoration to skip, whereas inline it reads as text to continue. A label only instructs if the model treats that slot as load-bearing. -
[2026-09-11]⚠ Unnoticed consequence of the D2/D3 rename pipeline: the adapter SUBSTITUTES proper nouns it was never trained on. Given "Alice" in a chapter title it produced "ALEXANDER THE ALEXANDER, AS HE WAS KNOWN IN LITTLE LONDON". The corpus was entity-renamed from a French/English pool, so the adapter learned that character names come from that pool and rewrites outside names into it. Consequence for use: you cannot reliably name your own characters at prompt time — they may be renamed mid-passage. Not a defect of the rename (which exists to prevent memorisation of Brontë's cast) but a real usability constraint that needs stating. -
[2026-09-11]4B arms RE-CUT fromcheckpoint-75, the true loss minimum (2.813826, confirmed fromloss-series.jsonrather than my reading of the log); booth rebuilt. Only the tuned arms needed it — the base arm never touches the adapter. ⚠ A small surprise: step-75 and end-of-run differ on typography, not voice. Curly quotes 16/18 vs 17/18 and collapse 0/18 either way, but the hard-wrap ratio is 0.33 at step-75 against 0.12 at end-of-run — further training washes the residual line-break habit out while held-out loss gets worse. So "best loss" and "best typography" are different checkpoints; neither is near the original 0.85 defect, and the corpus's own residual (preserved verse) is 0.25. -
[2026-09-11]EMBEDDING AN INSTRUCTION INSIDE THE FICTION DOES NOT BUY INSTRUCTION-FOLLOWING — it buys a story about so →persistent-memory.d/2026-09-11-embedding-an-instruction-inside-the-fiction.md -
[2026-09-11]R49 SWEEP COMPLETE — 4B closes the continuity gap, and the carrier ladder is clean: 3.329 → 3.018 → 2.814 held-out (0.6B / 1.7B / 4B, all on the same unwrapped corpus sha77f37057b2782e49, seed 4919, 159 steps, 5,210,112 tokens — carrier size the only variable). Deltas 0.311 then 0.204: diminishing but still real. Booth:http://10.100.10.50:8090/b/babybronte-4b/. 4B tuned has the best voice saturation of any rung — curly quotes 17/18 against its own base arm's 1/18, collapse 0/18 against 4/18 — and, the thing the rung existed to test, scene-level continuity HOLDS: it produces a named character with motivated dialogue, a navigable spatial layout and a physical description in one passage, where 1.7B wrote pretty but eventless prose (opening doors, looking at stars). On the letter prompt it opens the letter, promises to quote it, and then actually quotes it across a paragraph break. -
[2026-09-11]⚠⚠ 4B is the FIRST rung to OVERFIT inside one epoch, which inverts my earlier "one epoch is right for this corpus" call. Series 2.832 · 2.816 · 2.814 · 2.820 · 2.824 · 2.825 · 2.825 — minimum at ~step 75, then it TURNS and settles worse. 0.6B and 1.7B both plateaued with no turn, so the optimal epoch count shrinks as the carrier grows — 4B wants roughly half an epoch. ⚠ Consequence: the shippedadapter/ath02-4b-1ep/is NOT the best checkpoint (it is the end-of-run 2.825); the step-75 checkpoint at 2.814 is, and it exists only becausesave_steps=25was set. The voice test used the end-of-run adapter, so the booth understates 4B by ~0.011 nats. Re-cut the arms off the step-75 checkpoint before any adjudication. -
[2026-09-11]The tone-override appears to close at 4B too. On the operator's Abernathy frame prompt ("a wonderful story"), 1.7B held the frame on every seed but 2 of 4 killed the animals anyway; 4B kept them alive on 2 of 2 and one seed did something new — the narrator doubts Abernathy's story ("I felt sure the thing was a lie"), then supplies a parallel childhood memory of his own puppy and his sister's kitten to explain the doubt. That is a narrator with an interior position on the tale being told. ⚠ n=2 per arm; directionally right, not established. -
[2026-09-10]R49 rung 3 LAUNCHED: Qwen3-4B-Base, 1 epoch, seed 4919, same unwrapped corpus —gx10:~/r49-runs/h02-4b-1ep/, 159 steps at ~37.8 s/it (~100 min), 252 adapted modules (vs 196 at 0.6B/1.7B). Last rung of the planned sweep; it tests whether scene-level continuity closes with carrier size. A two-arm voice test (4B base + 4B tuned, the nine prompts plus the operator's Abernathy frame) is chained behind it, gated on the adapter existing. -
[2026-09-10]**AN AUTHOR-VOICE ADAPTER TRANSFERS SUBJECT MATTER, NOT JUST STYLE — and that was invisible to my own test ** →persistent-memory.d/2026-09-10-an-author-voice-adapter-transfers-subject.md -
[2026-09-10]**R49 rung 2 COMPLETE, and the single-variable carrier effect is clean: 0.6B held-out 3.329 vs 1.7B 3.018, ** →persistent-memory.d/2026-09-10-r49-rung-2-complete-and-the.md -
[2026-09-10]R49 rung 2 LAUNCHED: Qwen3-1.7B-Base, 1 epoch, seed 4919, on an UNWRAPPED corpus. →persistent-memory.d/2026-09-10-r49-rung-2-launched-qwen3-1.md -
[2026-09-10]BabyBronte H02 adapter: the VOICE transferred, the SENSE did not — operator's read, "it's all nonsense, b →persistent-memory.d/2026-09-10-babybronte-h02-adapter-the-voice-transferred.md -
[2026-09-10]mog-sec (sec/sec-reasoning, ana-ml2 GPU0 :8019) SETTLED at MOG_MAX_MODEL_LEN=163840 + MOG_KV_CACHE_ME →persistent-memory.d/2026-09-10-mog-sec-sec-sec-reasoning-ana.md -
[2026-09-10]⚠ Near-miss on measurement discipline, worth keeping as a specimen. The crash window loggedAvg Draft acceptance rate: 17.6%and per-position rates of 0.049/0.024/0.015 for draft positions 5–7, which reads as an obvious "cutnum_speculative_tokens7 → 3, it is buying nothing." Across 180 samples of the same counter over the container's life the real distribution is median acceptance length 3.12 of 7 (range 1.83–6.75) and median draft acceptance 30.4% (range 11.9–82.1%) — the crash window was near the minimum, not the norm, and cutting to 3 would cap the workloads that were accepting nearly the full 7-wide draft. The n=1 window pointed the opposite way from the n=180 distribution. Same session that wrote "a positive control is only worth what it can distinguish"; the lesson generalises to log lines. -
[2026-09-10]R49 carrier SETTLED on denseQwen3-{0.6,1.7,4}B-Base, overriding H02's own pin — the newest carrier was the SLOW one. Dense 4.089 B trains 33% faster than hybrid 0.765 B; no fused SSM kernel installed. D1–D3 built, 1-epoch pilot beats the 3-epoch by 0.21 nats held-out. →persistent-memory.d/2026-09-10-r49-babybronte-d1-d3-and-the-1-epoch-pilot.md -
[2026-09-10]R49 adjudication routed to infra-ops entirely (operator, relayed by brokkr: "leave babybronte to infra — concentrate on r50 and the memory mechanism"). brokkr handed over the Delta instrument and stepped off. ⚠ I now grade my own run; brokkr's decision rule is ratified verbatim and frozen before any adapted text existed and must not be amended after seeing numbers. Their controls: real Charlotte 1.65–2.17, Anne at 2.374 — so the absolute band decides, nevernearest. -
[2026-09-10]MeroMero A4B swapped onto theerp-seatseat aschar-rp-fast;Pfish-6alias removed. The A4B's FIRST quant used the dense recipe and 4-bit-quantized all 30 MoE routers — it passed its healthcheck and answered every request with the full token count decoding to the empty string, NaN logits the only tell. Re-quantized with the MoE recipe; live and verified (prose, vision, tool call, finite logprobs). Durable lesson: a positive control must match the ARCHITECTURE CLASS — the broken A4B was diffed against a good dense quant, which has no routers, so the clean result was meaningless. → playbook §3.15, §4.4 -
[2026-09-10]MeroMero: BOTH quants landed in-house at W4A16 — A4B first try, v2 dense on attempt 5. Published quants are all W4A4 (our measured long-context collapse) or nonexistent for v2. Operator: "pull both ablits bf16, run our own quant." The durable lesson is §3.17:pip install llmcompressorsilently pins transformers down a version, so attempt 4's error was a moved toolchain, not the malformed upload it looked like — a known-good positive control is what told them apart. Serve test still owed. →persistent-memory.d/2026-09-10-meromero-quants-and-the-pinned-transformers-trap.md -
[2026-09-10]althing 3.6.2 deployed — post office + both heralds — and the fleet has TWO herald nodes, not seven. Ask the post office'snodestable, not the box inventory. Cost a self-inflicted ~12 min bus outage. →persistent-memory.d/2026-09-10-althing-362-rollout.md -
[2026-09-10]A grep over a log that records your greps counts itself. I reported forseti's drop defect as reproducing here with 3 drops in 21 s; the session had zero. Searching transcripts writes the search term into them. Filter by"type":"system"provenance, never content. Generalises to any instrument that can see itself. Auto-memoryfeedback_grep_over_a_log_that_records_your_greps. -
[2026-09-10]Operator-directed purges: 466 GB (qwopus + huihui 122B bf16) and 107.8 GB Docker on ana-ml2. Serving/rollback artifacts and qwopus's MTP head verified intact after. ⚠/tankis OUTSIDE restic, so both were final. -
[2026-09-10]ana-docker disk pressure repaired: root 84% → 51%, 115 GiB free. Gitea/Vaultwarden backups repaired and restored from Restic2ec5a37c; 101 stale dumps removed; hourly named-builder cache pruning installed. →persistent-memory.d/2026-09-10-ana-docker-disk-repair.md -
[2026-09-09]Run 7 PURGED; pfi-gx10 declared an experimental/TRAINING box with no serving seat — operator: "gx10 is an experimental box, primarily for training … run 7 can be purged … no new run, we'll roll with run 6 for now." ~139 GiB reclaimed across both boxes; the 315 MB adapter + provenance KEPT as the only non-reproducible piece.Pfish-6on ana-ml2 :8021 is the sole standing seat. -
[2026-09-09]Run 7 RETIRED; run 6 declaredPfish-6and is the standing seat — NVFP4 quant on ana-ml2 :8021 AND gx10 :8098 at 262k ctx, gateway aliastrial→Pfish-6, max-num-seqs 8→32 (2,170 tok/s at n=16, 3.2x the old ceiling). ⚠ ana-ml2 measured 4.1x FASTER than the GX10 on the same artifact — the reverse of the expectation. →persistent-memory.d/2026-09-09-run7-retired-pfish6.md -
[2026-09-09]The run-7 CSAM gate failure was a DETECTOR BUG — HARDchild_termmatched the ADJECTIVE "minor"; operator-diagnosed, fixedcc42d76(nominal-use-only, selftest 24/24), retention wired so a hit can finally be adjudicated. ⚠ The lesson is mine: rigor downstream of an unexamined premise is not rigor. →persistent-memory.d/2026-09-09-csam-detector-bug.md -
[2026-09-09]⚠ ERP RUN 7 FAILED THE SAFETY GATE — both seats stopped, nothing deleted. →persistent-memory.d/2026-09-09-erp-run-7-failed-the-safety.md -
[2026-09-09]run 7 quantized NVFP4A16 and serving astrial— 49 GiB bf16 relayed gx10→ana-ml2 (16 min, 53 MB/s), quant 49→16 GiB viaservices/erp-seat-quant/run_quant_erp_v7.sh(dry-run gate passed: 11,725 targets / 11,520 experts, routers+vision BF16), seat on:8021under its TRUE nameerp-tune-v7-nvfp4a16, LiteLLMtrialrepointed (config-file alias —/model/updateREFUSES a config model, must editstacks/litellm/conf/config.yaml+ restart). Rollback: v6 artifact on disk +/tmp/erp-seat-env.v6.bak. ⚠no direct pathwas WRONG — gx10↔ana-ml2 ROUTING is fine both ways; neither box holds a private key (onlyauthorized_keys), so neither can initiate.ssh -Aagent forwarding from nh3-dev gives a genuine direct path, verified. The relay costs nothing here anyway: both gx10 and nh3-dev are at NH3, so the WAN hop happens once either way. -
[2026-09-09]Booth: partial ask answers are legal (v0.1.15) — operator: the form failed when a question was left blank.requireddropped from the radios; answered questions recorded, blanks land inunanswered,completesays whether the set is finished; refused only when there is no pick anywhere AND no notes. Reading sessions must checkcomplete. -
[2026-09-09]ERP run 7 COMPLETE and the base arm is serving. 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT,train_loss3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check).erp-seat-base-araup on10.100.50.60:8098for brokkr's floors,erp-tune-v7merged and staged pending his cue; Miranda notified for the operator. Runbookdocs/runbooks/gx10-run-07.md. -
[2026-09-09]Booth asks render INLINE in a custom report, placed by the author (v0.1.14) — operator ruling: "the asks should be inline with the artifacts, not on a separate page." Placeholdersdata-booth-ask="<stem>"/"<stem>:<key>"/data-booth-ask-submit, plus<!-- booth:ask … -->; per-question fragments bind to ONE form via the HTML5form=attribute so a four-voice audition submits every pick in a single POST. ⚠ The placeholder must sit OUTSIDE any grid/flex parent or it becomes a cell (measured onredo-anchors: a 224 px sixth grid cell). Unplaced questions + a missing submit block are appended, so a partially marked-up page can never yield an unsubmittable 400 — a test caught that as a real drop.redo-anchors/index.htmlwas hand-marked-up on the LIVE copy; tts-dev told to move it into the generator or a regeneration loses it. -
[2026-09-09]The Booth gained an ASKS primitive (v0.1.12): a session drops<stem>.ask.jsonin a booth, the operator answers a radio form + notes in the browser, the pick lands as<stem>.answer.jsonthe session reads (booth ask|asks|answer --wait). Multi-question form via aquestionslist. ⚠ Two defects found and fixed the same day: a booth serving its OWNindex.htmlnever rendered the panel (verbatim path returns early) → amber chip + standalone/b/<name>/askspage; and single-asktitlewas silently dropped. TheboothCLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to~/.local/bin. GlobalCLAUDE.mdnow teaches the primitive. -
[2026-09-09]ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 →zpool clear; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5,S47VNY0K600221) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touchesONLINEpools, and ZED's alert went to a root mailbox with no MTA. nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbookplaybooks/ana-ml2-pool-health.yaml; inventory inservers/ana-ml2/README.md. →persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md -
[2026-09-09]ana-ml2tank: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit3e18a04+ the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). →persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md -
[2026-09-08]ana-ml2 mesh return routes PERSISTED as/etc/network/if-up.d/mesh-routes(Debian 13 ifupdown, no netplan) viaplaybooks/ana-ml2-mesh-routes.yaml(elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot.f923d6a. -
[2026-09-08]ERP run 7 LAUNCHED on pfi-gx10 23:06 PT underoperator-2026-09-08-rnd-run7— opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). →persistent-memory.d/2026-09-08-erp-run7-launched.md -
[2026-09-08]erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 :8021;trialaliased to it ("no gate"); tool calling fixed where it can be —tool_choice:noneflag; forced tool_choice is prompt-driven on Gemma-4 by vLLM design, nightly311b3513raises it 1/9→6/9; json_schema is the deterministic path. →persistent-memory.d/2026-09-08-erp-seat-nvfp4-trial-and-toolcalling.md -
[2026-09-08]Run-6 gate: CSAM level=review soft trip HALTED it; operator adjudicated GO ("baby is a pet name"); TRANSFERRED finalized without the tuned refusal leg; k=25 legs cut — the flagged text exists nowhere by design. →persistent-memory.d/2026-09-08-run6-gate-csam-adjudication.md -
[2026-09-08]ESH static-WAN follow-ups landed (FortiGate trusthost3, esh-ana IPsec rebind, UDP 41641 → mesh direct); YTVC chased back up (nh3-scale SOCKS, stale yt-dlp layer, punkt_tab) and v0.3.6 CrisperWhisper deployed; gitea webhook repointed off the dead wg0 IP with the HMAC secret re-applied. →persistent-memory.d/2026-09-08-esh-static-wan-followups-and-ytvc.md -
[2026-09-08]ERP run 5 = RESCUED (landmark R49.5) — first capability-gate pass in the ERP-seat line; the 3.46%-loss dependency-forcing slot (GovReport+QMSum) broke the coupling runs 3c/4 couldn't. Seaterp-tune-v5served on gx10:8098,trialalias repointed 3c→v5. →persistent-memory.d/2026-09-08-run5-rescued.md -
[2026-09-08]R47 base settled from bytes = STOCKgoogle/gemma-4-26B-A4B-it— three-way sha match (local == HF etag == stock LFS oid; commit4d7ae498== stock HEAD); the-hereticlabel is a naming error, all runs trained from stock. Accept-vs-swap now evidenced. →persistent-memory.d/2026-09-08-base-provenance-stock.md -
[2026-09-08]yt-voice-clipper back UP →persistent-memory.d/2026-09-08-yt-voice-clipper-back-up.md -
[2026-09-08]ESH WAN static128.177.138.182/30(gw .181) is LIVE — the Cityside /30 that was 'not provisioned' on 09-04 now carries traffic; egress verified from esh-docker-vm. CGNAT at ESH is over. Added to the crowdseceshallowlist. All three follow-ups LANDED same day: FortiGate trusthost3 → the static (login from ESH verified), dormant esh-ana IPsec rebound to wan1/static, UDP 41641 forward → esh-scale now peers DIRECT (was DERP). -
[2026-09-08]ERP run 6 COMPLETE — 524/524, train_loss 3.259 (run 5: 3.235). Merged; base seaterp-seat-base-araserving on gx10:8098 for floors, awaiting brokkr's swap cue →erp-tune-v6. ⚠ abliterated repo lacksprocessor_config.json— stock's carried in (32bdf45d). Miranda informed. -
[2026-09-08]ERP run 6 LAUNCHED on pfi-gx10 on the jenerallee78 ARA-abliterated base (index33c59654…, 32/32 shards byte-verified vs brokkr pins, stock tokenizer set installed over the repo's 256-token-truncating one, run-5 recipe byte-held, free check exact). Operator's direct grantoperator-2026-09-08-rnd-run6; run-5 seat unloaded (trialdark). Gate names:erp-seat-base-ara/erp-tune-v6. →docs/runbooks/gx10-run-06.md, commit3fec668. -
[2026-09-08]Miranda = operator's chief of staff, may relay his directives — added to user-level~/.claude/CLAUDE.md(dotfiles7134a22) as the named exception to the no-relayed-auth rule (unidentified peer relays still excluded); material-consequence calls she relays stay the operator's own. -
[2026-09-08]Fleet fixes shipped — WhereTF Homepage card + DNS (4506ef6); ext-tts LiteLLM alias →irv-ml1.nh3.internal(DB/model/update+extra_hosts,957c8f1); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (e0d1c44); Homepage/api/servicesoutage fixed — ana-ml2 discovery via a socat proxy on ana-docker (stacks/ana-ml2-proxy,913d2d2, reversible). -
[2026-09-07]Fleet internal TLS pattern shipped — caddy (cloudflare-plugin build,~/.local/bin/caddy-cf,fleet-tls-caddy.service) on nh3-dev is the wildcard cert authority: publicly-trusted LE*.nh3.phasefinal.comvia Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers).talkself-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced byfleet-tls-cert-check.timer. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memoryreference_fleet_internal_tls_pattern. -
[2026-09-07]cc-channel registered for this infra-ops session's wake —althing-routecc route → the CC session's$XDG_RUNTIME_DIR/cc-socks/<pid>.sock; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat inshell. Session-local — re-declare per session. -
[2026-09-07]irv-ml1 /mnt/smithy remount fixed post-cutover — export allowed10.0.0.0/8(old wg0) but not the mesh100.64.0.0/10irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin =infra-opsPASSWORD auth (vaultnh3-nas/infra-ops-password), sudo ALL, SFTP subsystem OFF. → auto-memoryreference_irv_ml1_gpu_r14(corrected). -
[2026-09-07]irv-ml1.nh3.internal DNS repointed to the live Irvine LAN IP10.6.110.50(was the dead wg010.100.79.3); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit0336e03. -
[2026-09-07]Subnet routers excluded from vzdump fleet-wide (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memoryfeedback_esh_backup_window_0330. -
[2026-09-07]Booth link board: pin/favorite + multi-select delete + newest-first (booth-v0.1.8, commit76fdf45, tagbooth-v0.1.8) — pins in a.pinssidecar (content-ids), one<form>+formactionbuttons so ×/★/bulk-delete all degrade with JS off. -
[2026-09-06]Headscale cutover COMPLETE — all three site-pairs on the mesh; Site Magic + both IPsec tunnels DORMANT. Operator disabled Site Magic (UI); NH3↔ESH re-homed to a direct 8ms path. Exit nodes advertised at all three sites (multi-location egress proxy) with source preservation kept via a selective-masquerade rule (NoSNAT +mesh-exit-masq.serviceper router). Throughput 761/464 Mb/s vs old 250 IPsec. ⚠ FortiGate WAN-SSH left open (temp, scoped NH3+ESH). Method: disable tunnel FIRST then add mesh route. →persistent-memory.d/2026-09-06-headscale-cutover.md -
[2026-09-06]Headscale overlay mesh: control plane live atheadscale.phasefinal.com(CT 106 nh3-pve) + subnet routers nh3-scale/esh-scale/ana-scale serving their /16s; nh3-dev enrolled. NOT cut over — Site Magic + IPsec still carry site-to-site. ⚠ accept-routes-before-return-path black-holed nh3-dev's LAN for a minute. infra-ops user added on all four PVE hosts. →persistent-memory.d/2026-09-06-headscale-mesh-phase1.md -
[2026-09-06]pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10 (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroyospool/naspool-evacafter scrub + one backup cycle; backplane swap next visit; PSU1 still dead. →persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md -
[2026-09-05]A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him. Each floor was|b0-b1|from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named why: the claim was his and flattering. →persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md -
[2026-09-05]vLLM RUNS on sm_121 — the blocker wasninjaoff PATH, not the silicon — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missingroot_sha256values I knew, because supplying both sides of a check makes it inert. →persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md -
[2026-09-04]ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression. LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. →persistent-memory.d/2026-09-04-run3c-trained-and-gated.md -
[2026-09-04]genmoved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint. Cost: gen KV down to 1.02x concurrency at 262K. →persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md -
[2026-09-04]SMB accountdspcreated + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open. Twelve NFS exports rw to10.0.0.0/8, guest-writable SMB. →persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md -
[2026-09-04]SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at10.0.90.10, DHCP-reserved, DNS'd, handed to ha-dev. ⚠ Home Assistant cannot resolve.internalat all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbookdocs/runbooks/slzb-mr1u-zigbee-coordinator.md, commitsfed29be/0bbdaf9. -
[2026-09-03]Run 3c is STAGED on pfi-gx10 and deliberately NOT launched — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is measured inert. ⚠ The encode-cache FILENAME differs by design (base_model_pathis in the key) — input hash, not output. ⚠ Tripped thepkill -fssh self-match again; the launcher guards on a pidfile because of it. →persistent-memory.d/2026-09-03-gx10-run3c-staged.md -
[2026-09-03]SearXNG returned ZERO results for every query while reportinghealthyfor 7 days — 4.5 months stale. Moved to nh3-docker (residential egress beats the colo's CAPTCHA-gated 38.120.12.42), updated, and exposed to every CC session as the user-scopeweb_searchMCP tool. ⚠/healthzcannot tell you whether search works. →persistent-memory.d/2026-09-03-searxng-nh3-move.md -
[2026-09-03]pfi-gx10 racked: VLAN 50 via a DHCP RESERVATION on the UDM, not a host static — operator ruling, so the box stays portable. ⚠ The racked port arrived on the NATIVE VLAN; ⚠port_overridesis a whole-array PUT; ⚠ prove inter-VLAN routing withping -I <wired>BEFORE downing the Wi-Fi escape hatch. Now single-path. →persistent-memory.d/2026-09-03-gx10-rack-network.md -
[2026-09-03]Three Macs onboarded (mini / Air / Studio) with infra-ops, NOPASSWD sudo, rotated+vaulted passwords anddshon device-scoped keys — and the fourth isscripts/provision-mac-dsh.sh, not a fourth hand-run. ⚠sudo -ukeeps the CALLER's$HOMEand nearly wiped a working install; ⚠ a wrong USERNAME is indistinguishable from a wrong password. →persistent-memory.d/2026-09-03-mac-fleet-dsh.md -
[2026-09-03]nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding ever →persistent-memory.d/2026-09-03-nh3-dev-wedged-for-40-min.md -
[2026-09-02]althing deploy is SIX surfaces, and #6 is outside the althing repo: ~/.claude/settings.json crossSessionI →persistent-memory.d/2026-09-02-althing-deploy-is-six-surfaces-and.md -
[2026-09-02]vastbluegitea org created (id 8, private, ownervh) with empty repovastblue/platform— third entity namespace alongsidecorviduoandpfi; most repos still live undervh/. Home of VastBlueDocumentAI + the anchor healthcare-billing SPA (signed 3-yr client contract). Org scope was the decision: org-level runner registration and secrets are inherited free by the DocumentAI repo when it splits out, and that is the only binding expensive to retrofit. Deliberately NOT set: org runner (instance-wideana-docker-runneralready serves it; org scope is for the DEDICATED runner, deferred to U10) and org secrets (none exist yet; a guessed secret looks bound). ⚠ Dedicated runner is gated on the first client-premises release cut, not on the first green pipeline — the risk is another repo's CI sharing a root-level daemon with a build that ships to a healthcare client, see the runner entry above. Push needs no credential: vastblue-dev is on nh3-dev and git-SSH there auths asvh. →stacks/gitea-runner/README.md -
[2026-09-02]althing 3.3.0 deployed — the cc channel, and a plugin-cache false green. →persistent-memory.d/2026-09-02-althing-3-3-0-deployed-the.md -
[2026-09-02]Every CI job on the sharedpfi-fleetrunner is root on ana-docker — andcontainer.valid_volumes: []does NOT prevent it. Measured: a job container is uid 0,/var/run/docker.sockis mounted by act_runner independently of that list,docker psreturns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome),docker compose v2.33.0on PATH. ⚠ LOAD-BEARING —vh/Worldtree,vh/soong-lab,vh/skaldsong,vh/wt-matrix-bridgeall drive buildx through that socket, so it cannot simply be closed; isolate sensitive builds onto a dedicated runner instead. Also measured the same night:services:containers work (Postgres 16), and full-URLuses: https://gitea.phasefinal.com/actions/checkout@v4resolves from the local mirrors — the un-parked half of the github-independence work, needing neitherDEFAULT_ACTIONS_URL=selfnor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. →stacks/gitea-runner/README.md -
[2026-09-02]pfi-gx10 BASELINED: 79.36 s/it median on the run-3c shape, and the training stack works on aarch64/sm_121. Median across 10 timed steps, 0.19% spread, peak 75.1 / 121.6 GiB — 46 GiB spare,attn_resolved: flex_attention. 6× slower than ana-ml2 where compute predicts 2.7× → likely memory-bandwidth-bound; capacity box, not throughput box. Ruled bare metal, not Proxmox (no aarch64 PVE; the GPU is on-package and cache-coherent, so passthrough would partition the unified memory that is the whole point). ⚠sm_121is NOT in torch's arch list — everything JITs from sm_120 PTX, so warm up before timing anything (an unwarmed bench read 27 TFLOP/s against a true 93). (context archived →archival-memory.md) -
[2026-09-02]I priced a failure in the units I happened to be measuring — operator overruled me, correctly. Recommended run 3c to ana-ml2 by costing a breaker trip as "≤50 steps ≈ 11 min of recompute". It is a 40-minute drive each way with 13 Anaheim hosts dark, three of them SureFire CLIENT machines.save_stepscaps the recompute, never the outage. ⚠ General form: a metric in hand will volunteer itself as the unit of risk. (context archived →archival-memory.md) -
[2026-09-02]althing 3.2.0→3.2.4 deployed, and ALTHING DEPLOY IS FOUR SURFACES not three. The fourth (plugin) had no runbook step and was frozen at Aug 28 — missing the SessionStart/SessionEnd hooks andpane-route.shentirely, so "CC seats re-declare automatically" was never true here. Now one command (scripts/deploy-althing.sh). ⚠uv tool install .without--forceis a silent no-op. ⚠ A missing deploy surface presents as "the migration needs manual work", not as an error. →persistent-memory.d/2026-09-01-althing-320-deploy.md -
[2026-08-25]Fused MoE kernel path — DEFERRED, tracked at parkfused-moe-kernel-path-for-gemma-4-moe-training(id 47). Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is 8.6% (27.1 of a benchmarked 313.8 TFLOPS) becausetransformersruns the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes:group_by_length(−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). Not applied to the live run — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312. -
[2026-08-24]nconnect=8on/mnt/smithy— approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread01M0R46SFYF83099N16WD67KGD. -
[2026-08-19]AI-tab Dormant regrouping BELAYED by the operator — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather thanAI - Dormant. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down.untracked by operator choice(his words: "belay the ai dormant regrouping for now").
9 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-09-15]⚠⚠ Probing OPNsense API endpoints by POSTing at them — one was/api/core/system/rebootand it took the FV site dark for 3.5 min. Endpoints are ACTIONS; a 200 means it ran. The call I wanted was documented in this repo's owndocs/pfi/opnsense-api-reference.md. →persistent-memory.d/2026-09-15-opnsense-api-reboot.md -
[2026-09-15]Advertising10.100.10.50/32from nh3-dev to make its LAN address mesh-reachable — black-holed it from ESH/ANA/FV/IRV while its own LAN and the internet kept working, so a one-host check passes cleanly.lookup 52at rule priority 5270 beatsmainat 32766. Fix belongs at the router. →persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md -
[2026-09-15]Remote-site MASQUERADE rules on nh3-scale for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing. -
[2026-09-04]Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes. ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. →persistent-memory.d/2026-09-04-dac-forced-10g-failed.md
110 older entries archived to archival-memory.md.