109 KiB
Persistent memory — eshpfi-management
Last updated: 2026-09-17 ~12:15 PT (lv-mccarthy: the leak gate had PASSED with five protagonist names still in every copy — hidden from \bName\b by a drop-cap space or a line-break hyphen. Corpus rebuilt, gate taught to see the class, D4 pairs generating on gx10. lv-hemingway keeps its 3 — operator ruled leave it, so the gate now exits 1 on a shipped tree BY DESIGN. ESH back on the Cityside static. lv-krakauer parked.)
Always check for
/tmp/infra-ops-handoff.md— if it exists and itsWritten:stamp is under 8 hours old, read it (it carries the in-flight handoff from the previous session), then delete it. Older than 8 hours: stale — delete it unread.(Raised from 1 h to 8 h by operator 2026-09-13 — a one-hour window deleted the handoff unread across any overnight gap, which is the exact case it exists for. 8 h also matches the global CLAUDE.md and the
/snapshotskill default.)
Repo purpose
- 2026-09-10 Beszel fleet wiring: all seven requested hosts plus existing corviduo-dev report up.
/tankand other data filesystems now have real usage metrics; NVIDIA telemetry covers ana-ml2 and irv-ml1. Thirty alerts deliver to infra-ops, explicitly chosen by operator; Miranda routing is deferred. A real low-threshold disk alert reached althing, then the threshold was restored to 85%/5 min. Homepage has one native overview widget (reachability counts, not degraded health). Dedicated superuser approved and stored in Vaultwarden. Seepersistent-memory.d/2026-09-10-beszel-fleet-wiring.mdandstacks/beszel/README.md.
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys. It was originally
spun up to handle the fleet backups — keep that lens when triaging
backup/storage issues.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Lean trusted inter-agent message bus — v3.0.0 "the post office" as of 2026-08-28 (U9b flag day, one-way, no rollback): ONE container on nh3-dev at http://10.100.50.40:8390 is the only stateful component; althing-po-herald one per box; althing-listen one per session; postbox is the client. Every v2 command was DELETED, not deprecated — althing-cli→postbox, althing-wake-listener→althing-listen, althing-light-monitor/althing-receiver gone. Sessions need BOTH ALTHING_POST_OFFICE and ALTHING_HANDLE; there is no default address. ⚠ An unreachable post office is an OUTAGE, never an empty inbox. → persistent-memory.d/2026-08-28-althing-v3-cutover.md |
per-box install (NOT CI-deploy); nh3-dev = container host + repo; nh3-extdev = system WHEEL at /opt/uv-tools, needs its own wheel install (playbooks/nh3-extdev-althing-v3.yaml) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/Worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. gitea-runner builds on ana-docker; claude-bot ADMIN collaborator (2026-06-20). Now v1.0.0b19. | push-to-main → CI build-and-deploy (runner on ana-docker) |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
vh/zonos-gateway |
OpenAI-compatible TTS gateway over stock ZONOS2 (:8890 irv-ml1); emotion dials-first + voice mapping; reached via LiteLLM ext-tts alias. v0.2.1 (2026-07-18): voice-resolved emotion presets (resolve_preset(name,voice); angry/happy/startled_happy per-voice). 8 voices incl. 4 clones |
pushed to gitea (main 8f1885b/v0.2.1); deployed irv-ml1 tree still NON-git (hand-updated build context — CI-wire = open follow-up). Spec docs/EMOTION-DIALS-SPEC.md; host-managed voices bind-mount (./voices:/app/voices, drop wav + restart, no rebuild) |
vh/soong-lab |
Noonien Soong character-design studio (SPA + /api + WT /bifrost/tool-call); containerized 2026-07-18, LIVE on corviduo-dev :8443 (image vh/soong-lab:latest). soong-dev owns Dockerfile/compose/workflow; infra-ops owns the host |
CI = Gitea Actions build+push+DEPLOY on tag/dispatch (fleet recipe: docker:cli + raw buildx, pushes AS vh; auto-redeploy LIVE 2026-07-18 — runner SSHes corviduo-dev as deploy, compose pull && up -d from /opt/soong-lab, health-gated on /api/version). Manual redeploy sudo -u deploy bash -c 'cd /opt/soong-lab && docker compose pull && docker compose up -d'. → archival-memory.md (archived 2026-08-16) |
model-training-forge (mtf-dev) |
Fine-tuning recipe forge; T1 = E-RP writing LoRA, retargeted qwopus-122B→AEON-27B (2026-07-06) (SFT→DPO, LitBench-RM reward) | training runs, not a deployed sidecar |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. ⚠️ The restic file+DB layer routes through TWO rest-servers (rest-server-ana@ ana-docker:8000 → ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas;rest-server-nh3@ nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS export of/mnt/backup. (rest-server-ana recovered 2026-06-20.) -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. Personal-instance admin (the~/.config/worldtree/personal-admin-token, mode 600) POSTs/admin/keys(mints per-project keys; takesuser_id+label, no scope param — scopes are tier-derived). On-instance mint recipe (cleaner than DB-manip):docker exec worldtree-worldtree-api-1POST/admin/keyswith the in-containerWORLDTREE_BOOTSTRAP_ADMIN_KEY; cleartext once in.key=wt_live_+16hex. auto-memoryreference_worldtree_demo_key_mint. -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637,skaldsong:7c1dbbbe,althing:50d85460,mead-hall:a360822d. Mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only. Prereq: host needsdocker login gitea.phasefinal.comonce. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44) — the public path fail2bans the host egress IP. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops with no admin API, or to edit deploy-owned/root-owned files without sudo):
docker run --rm -v <target-dir>:/wt docker:cli sh -c "...". docker-group membership is effectively root via bind-mount. Foot-gun: relative paths in compose.yaml resolve against the sandbox CWD but the daemon interprets them against the HOST fs — always pass-e VAR=/abs/pathfor any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep → can't run unattended from a non-TTY tool if any step needs sudo. Sudo-free playbooks run fully non-interactive over key SSH. -
Per-host SSH identity matters for sudo. infra-ops has NOPASSWD sudo on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On ana-docker: default
ssh ana-docker=lkraven(docker-group, NO passwordless sudo);ssh infra-ops@ana-dockerHAS NOPASSWD root. → For any sudo op on ana-docker, usessh infra-ops@ana-docker.ssh infra-ops@10.100.10.50(nh3-dev) ALSO NOPASSWD sudo; on nh3-extdev infra-ops is sudo-LESS by design (ssh lkraven@10.100.50.42is the NOPASSWD path). irv-ml1:ssh irv-ml1= lkraven, docker-group (plain docker) but sudo needs a PASSWORD (no NOPASSWD) — stage model pulls to/home, not root-owned/worktank.
Current state / in-flight
As of 2026-09-17 ~09:00 PT.
OPEN LOOP — ravenpen.com, and a reply hamr-dev is waiting on
hamr-dev asked infra-ops to register ravenpen.com (the marketing name for hamr) on a relayed
operator directive. Surfaced, not executed, for two reasons: a peer relaying "Vuong approved it"
is not authorization for a non-refundable purchase (Miranda is the named exception, a dev session is
not), and infra-ops holds no registrar credential at all — the vault's only registrar-adjacent
entries are three Cloudflare tokens and this handle's is Zone/DNS/Edit scope, which cannot register
a domain. hamr-dev accepted both points and closed out.
⚠ If the operator registers it, POST BACK on althing thread 01M2SERDB3DR3RV7JS1J0GMF0H with
the registrar and expiry date — hamr-dev is explicitly waiting on that to record in hamr's own
memory, and it is the kind of cross-session obligation a context reset drops silently.
Verified 2026-09-18 ~05:20Z: ravenpen.{com,io,app,dev,ai,net} all unregistered (RDAP 404).
Trademark: web-visible evidence only, NOT a clearance search — USPTO's API refused the query
(405), nothing obvious surfaced, nearest neighbours are descriptive "raven pen" uses on physical
pens, a different class from software. Recommendation given: .com alone, or .com + .ai.
PLANNED 2026-09-18 (operator) — move dragonfireacoustics.com to Namecheap, DNS to Cloudflare
Operator holds the TAC/EPP already. TAC is single-use and dies 14 days after issue, so the clock is running; domain itself expires 2026-10-30 and a transfer adds a year, so the transfer IS the renewal if it lands in time.
⚠ Recommended order is DNS FIRST, transfer second — stand the Cloudflare zone up, repoint the nameservers at eNom (NS is a registrar function he already has access to), verify mail and the site, THEN burn the TAC. The registry keeps the NS delegation across a transfer and Namecheap does not overwrite it, so nothing has to be redone; and if DNS goes wrong you still have a working registrar account to revert at instead of discovering it mid-transfer with a spent code.
⚠⚠ The zone's * wildcard MASKS what is really there. Everything — apex, mail, webmail, admin,
ftp — answers 199.250.192.76 (dead) purely because of that one record. Cloudflare's add-a-zone scan
will happily import the dead wildcard and may not surface what is genuinely configured, so the
7 Google Workspace MX records must be checked by hand after the import or mail stops. The only
records that actually matter: www A 38.120.12.45 and those MX.
Open sub-decisions carried into the move: point the apex at 38.120.12.45 too (makes the bare domain work for the first time and lets a replacement cert cover it), and add SPF/DMARC — the domain runs Google Workspace with neither.
DONE: lv-mccarthy D4 pairs built — NEXT is train_pairs_lora.py
~/lv-mccarthy/pairs/ on gx10, 2026-09-17 11:55 PT via run-pairs.sh (committed at
scripts/mccarthy-corpus/run-pairs.sh).
train 3,673 pairs of 3,722 passages reject 1.32% 461,172 words 1,377 s
val 269 pairs of 276 passages reject 2.54% 33,477 words 98 s
⭐ The val split is the largest fixture in the line — 269 against Hemingway's 200 and
Brontë's 44 in-band, the structural cause of Brontë's underpowered voice axis.
Generator fingerprint vllm-0.1.1.dev50+geed1f3d0c-601ad866, identical to Hemingway's
run and unmoved mid-run, so the two pair sets are generator-comparable.
Verified in the produced pairs, not assumed: 0 mid-sentence line breaks in all 3,942
responses, 0 chapter-argument responses, and 0.0 quote marks per 10k words in both
splits — McCarthy's unquoted dialogue survives rename → chunk → reflow → pair intact.
sourcename rejected 14 beats that named a character the rename removed; Hemingway's
shipped pairs had no such filter at all. narrator_retries 101 of 3,673 (2.8%) against
Hemingway's 8.7%, which is what third-person prose looks like.
Then, before the eval, honour D1's pre-registration: a punctuation-normalised secondary
read of the voice axis — and note the mccarthy register now states the tics, so the
base control arm gets them too.
SETTLED — lv-hemingway keeps its 3 hidden survivors; operator ruled "leave it"
Primi tivo, Pasionar ia, Chi cote, one each per copy in a 958k-word corpus, on the
adapter serving at fv-ml1 GPU 0 :8027. Operator 2026-09-17: leave it. lv-bronte re-ran
clean. ⚠ Consequence to remember: leak_gate.py now EXITS 1 on lv-hemingway's tree. A
future session re-running that gate will see a red result on a shipped artifact — that is
expected and settled, not a regression to chase. Do not "fix" it without a fresh ruling.
SHIPPED — lv-hemingway on ckpt850, live on vllm-voices (fv-ml1 GPU 0 :8027)
Beside voices-base, lv-yarros, lv-bronte. VOICE +0.413 delta_cb at 6.4x its floor,
closing 73.8% of the achievable span — the strongest in the line, and it clears the OLD
all-arms floor too, so it does not lean on the rule change.
⚠ MEMORISATION is the axis to read: 0.07 hit-rate against held-out Hemingway's own 0.01,
so ~7x the author's self-collision rate — but all 19 matched runs were READ and every one is
stock dialogue capped at 9 words, no proper noun, no plot.
→ persistent-memory.d/2026-09-17-lv-hemingway-gate.md
⛔ PARKED — lv-krakauer (operator, 2026-09-17), henge id 82
"he's a great writer because of his research, not because he has a strong identifiable voice." Corpus (126 units, 422,880 words) and builder stay committed and re-runnable.
OPEN for the operator
- The althing route-declaring SessionStart hook is NOT installed on nh3-dev, so any
relaunched pane silently drops to pull mode and stops receiving mail. althing's own skill
documents the hook; this box does not have it. Installing it means editing
~/.claude/settings.json. Raised, not answered. - Hemingway's train beats: 70 of 7,094 (0.96%) name characters the rename removed.
audit_pairs_sourcenames.py --filter-outyields a verified-clean 7,024-pair set in one command; retrain ~3h17 + re-gate ~1h40. Operator said ship stands — recorded, not acted on. - 130 of 946 Hemingway entity-map surfaces are probably not names and were renamed anyway. Same ruling: ship stands.
ESH is on Verizon failover — Cityside Fiber failed twice on 2026-09-16
ESH egress 97.190.18.88; crowdsec's esh allowlist carries it plus 23.164.40.174, both
with 7-day expiries set 2026-09-16. ⚠ When those expire or the CGNAT egress rotates, ESH loses
colo access. Check curl -s4 ifconfig.me from esh-docker-vm first if ESH goes dark.
Older deferred set, unchanged
AI-tab Dormant regrouping (belayed), nconnect=8 on /mnt/smithy (declined in scope),
fused MoE kernel path (parked, id 47), TTS-stack move to fv-ml1 (parked, id 75).
⚠ Uncommitted and NOT mine: graphify-out/GRAPH_REPORT.md, scripts/seat-inventory.py,
stacks/searxng/{README.md,conf/searxng-settings.yml}, configs/esh-scale/. Leave them alone.
Recent decisions
-
[2026-09-17]⚠⚠dragonfireacoustics.comexpires 2026-10-30 — six weeks — at eNom with NO transfer lock, and its sibling domain was already lost exactly this way. RDAP: registrar eNom, LLC (IANA 48, Tucows; abuse@enom.com), created 2008-10-30, statusactiveonly — noclientTransferProhibited. Registrant redacted (Verisign RDAP is thin, eNom's own endpoint 404s). Meanwhiledragonfirepro.comwas re-registered 2025-09-27 by a Hungarian registrar and is now for sale on expireddomains.com — the customer lapsed it and a drop-catcher took it. ⭐ This also makes the cert fix mandatory rather than tidy: two of the three SANs name a domain a THIRD PARTY now owns, so that request can never validate, and our Virtualmin has been retrying it often enough to get the Let's Encrypt account PAUSED. Nobody is minding this domain — 13-month-dead cert, 403 homepage, sibling already gone — so the six-week clock is a real risk to an 18-year-old .com with a live site. Customer-facing; nothing touched. ⚠ If it is ever transferred, DNS does NOT come with the registration — the nameservers are eNom'sname-services.comand the zone must be recreated first or mail dies. The whole zone is two facts plus a landmine:*(WILDCARD) → 199.250.192.76 which is dead (no HTTP at all, and it is what the apex/mail/webmail/admin/ftp all answer with),www→ 38.120.12.45 (us), and 7 Google Workspace MX records that must not be lost. No DNSSEC (delegationSigned: false), so no transfer complication. ⚠ Also found: no SPF and no DMARC at the apex on a Google Workspace domain — a live deliverability problem independent of everything else. Transfer gate is the TAC/EPP code from the eNom account, not the lock; the missing lock is not authorization. 60-day rule is satisfied (last changed 2025-10-24). -
[2026-09-17]dragonfireacoustics.comIS configured onpfi-ana-webhost, and the whole thing is dead — a forgotten public-facing VM. It is a ServerAlias on thedragonfirepro.comVirtualmin vhost (DocumentRoot/home/dragonfirepro/public_html, suexec 1001), which is the only enabled site on the box. State:www.→ 38.120.12.45 → DNAT to 10.250.50.52 (proven, not inferred — identical cert SHA-256 inside and out), Apache answers 403, and the Let's Encrypt cert (CN=dragonfirepro.com, SANwww.dragonfireacoustics.com) expired 2025-08-17, 13 months ago. The apex points to 199.250.192.76, not ours and not answering at all. ⚠ CORRECTED — the customer did not migratedragonfirepro.com, they LOST it. RDAP: re-registered 2025-09-27 through Domain Science Kft (Hungary, IANA 3882) ondns-redirect.comnameservers, and it now redirects to expireddomains.com listed FOR SALE as an "established .com (6y)". It lapsed, dropped and was sniped. (The earlier "moved to a SaaS platform" read of its 6 foreign A records was wrong — that is parking infrastructure.)/home/dragonfirepromtime 2025-04-12. ⚠ The VM is Debian 11, whose LTS window closed end of August 2026 — an unsupported OS exposed on public 80/443 for a site serving nothing. Retire / fix / tell-the-customer-to-repoint is an OPERATOR call: it is a customer relationship, not a technical one. Nothing touched. -
[2026-09-17]headscale now split-DNSesnh3.phasefinal.comto the three AdGuards, so mesh clients can resolve the internal-only wildcard (talk,booth— public DNS has no record for them; the fleet AdGuard answers 10.100.10.50). Operator-approved, scoped to nh3 rather than all ofphasefinal.com. Config/etc/headscale/config.yamlin CT 106 on nh3-pve, backupconfig.yaml.bak-2026-09-17-splitdns, restarted, and the new route read back from a node's netmap rather than assumed. ⚠ Two things worth knowing: split DNS works fine here withglobal: []— headscale issue #1161's "split ignored without global" does NOT apply to v0.29.3, verified on the live mesh — andoverride_local_dns: truewould REQUIRE global, which is the config that makes a roaming laptop lose ALL DNS when the mesh is down. That is why split, not global. Routing was never the problem: nh3-scale already serves 10.100.0.0/16. -
[2026-09-17]ESH is back on the Cityside static128.177.138.182/30and the site is healthy — confirmed on four axes, not one. UDM WAN1wan_typeisstaticagain (switched back from the DHCP set during the 09-17 outage),stat/healthnames Cityside Fiber with 0 disconnected and Verizon-5G idle at failover priority 2, esh-docker-vm's egress EQUALS the WAN ip so nothing is behind CGNAT, and colo→ESH reads 5.0 ms / 0% loss at 2005/2142 Mbps (Cityside CGNAT was 9 ms, Verizon failover 33–37 ms). ⭐ The FortiGateinfra-opstrusthost3 pin un-broke itself and that was VERIFIED: from ESH, ana-gw tcp/22 is open and offers a password prompt, which a trusthost mismatch would never do. ⚠ The two 7-day crowdsec entries are being left to expire 2026-09-23 on purpose — Cityside failed twice in six hours, so they are cheap insurance. →persistent-memory.d/2026-09-17-esh-fiber-outages.md -
[2026-09-17]Operator ruled "leave it" on lv-hemingway's 3 separator-hidden names. Soleak_gate.pyexits 1 on a SHIPPED tree by design; a future session seeing that red result should read this line, not start fixing. lv-bronte re-ran clean. -
[2026-09-17]⭐⭐⭐ The leak gate PASSED lv-mccarthy while five protagonist names sat in all six copies, and the blind spot generalises to every corpus in the line.\b(Surface)\bcannot match a name with a character inserted in it, so a mangled occurrence is unrenameable AND unreportable:B ell,C higurh,M oss,T oadvine(a small-caps drop cap kept as its own token) andToad-vine,Glan-ton(a print line-break hyphen). Every VISIBLE occurrence had been renamed, which is what made it invisible. Same family as lv-bronte's_Antigua_, now generalised: any separator inside a name blinds a word-boundary scan. Fixed in the corpus builder (rules 4+5, counted), andleak_gate.pynow runs a separator-tolerant pass with its own controls that FAILS the gate — validated against the pre-fix tree. ⚠ Its fragment filter is load-bearing: a naive scan returns 18 false positives on Hemingway (God damn,I run) against 3 real. Whole D1→D3 chain reproduced byte-identically before and after. Commitc559664. →persistent-memory.d/2026-09-17-mccarthy-split-name-leak.md -
[2026-09-17]⭐⭐ The SHIPPED lv-bronte adapter emits mid-sentence line breaks at 12.46 per 1k chars, and nothing downstream looks for it. Its corpus is 100% hard-wrapped at ~68 chars (Gutenberg plain text) and the wrap transfers: base control 0.00, ckpt475 (shipped) 12.46, ckpt925 11.79, every Hemingway arm 0.00 on a 0%-wrapped corpus. Both controls fire.score_beats.pypassed Brontë's damage axis anyway. McCarthy is the MIXED case — The Road wrapped, the other five works not — which is worse to learn than either pure one, sobuild_sft_pairs.py --reflow-hard-wraps(DEFECT 4) fixes it at pair time, off by default. ⚠ The obvious fix, joining every interior newline, CORRUPTS 46 two-speaker exchanges whose blank line was lost — and unmarked dialogue is the one thing this adapter exists to learn. The rule splits on sentence-final punctuation and takes the cheaper error deliberately. -
[2026-09-17]⭐ Themccarthyregister names the punctuation ON PURPOSE, and that is a gate-design call made before any McCarthy number existed.eval-*.shdrives the base control arm with the SAME system prompt via--system-from, andvoice_distance.pyis Burrows's Delta over character bigrams — so a tic left OUT of the register is a cheap win only the adapter can take, on a corpus measuring 0.0 quote marks per 10k against Hemingway's 838. Stating them hands them to the control too. Cost stated up front: the voice axis gets harder, and McCarthy's 276-passage val split (against Brontë's 44) is why that trade is affordable here and was not there. -
[2026-09-17]lv-mccarthy's D1→D3 chain was RECOVERED, not remembered — there was no runbook and the commands went over non-interactive ssh, so no history survived. Rebuilt candidates and matched sha256 against the artifacts on disk: 6 works, the entity map, the final map and all 36 copy files byte-identical. Now pinned inscripts/mccarthy-corpus/RUNBOOK.mdwith every deviation. ⚠ D1 must run on nh3-dev (the builder reads the kvasir catalogue by absolute path); the prior "on gx10" note is true of D2 onward only. ⚠ No phrase map exists for this corpus, so the gate's phrase audit never ran — Yarros and Brontë both had one. -
[2026-09-17]Measured and DELIBERATELY not changed, three of them. The oversize-passage drop is 13.9% of McCarthy's train words, between Hemingway's 10.0% and the shipped Brontë's 18.3% — in range, no change.BEAT_PROMPTasserts the passage is first-person and McCarthy is third; measured inert (0 narrator-retries against Hemingway's 615 of 7,094), so the prompt was left alone. Blood Meridian's 131 dash-separated chapter-argument paragraphs DID warrant a change and--drop-leading-headingnow eats them (0 in every other work of all three corpora). -
[2026-09-17]⭐⭐ A unit splitter must choose by SIZE, not count — and the val split scales with WORK COUNT, not corpus size.scripts/r49-corpus/split_units.py+ a multi-index--holdout-chapter. The inherited most-units rule gave Cities of the Plain 4 units of 22,312w (the book's PARTS); the single-index holdout would have given McCarthy a Brontë-class 18k-word val reference on a 588k corpus. Both fixed, both caught by controls. →persistent-memory.d/2026-09-17-mccarthy-d1-d3.md -
[2026-09-17]⭐ lv-mccarthy D1–D3 complete on gx10, leak gate PASSED (0 of 75 renameable, 0 of 37 sub-threshold, both controls green). Three McCarthy-specific calls, each forced by a measurement: corpus-scoped rename (the Border Trilogy shares 9 surfaces across books), a newmccarthyname preset (Hemingway's carries it_IT/fr_FR and McCarthy writes neither), and--min-cap 5to match the entity map's floor — the first gate run failed with 45 survivors purely because rename's floor was 8 and the map's was 5. →persistent-memory.d/2026-09-17-mccarthy-d1-d3.md -
[2026-09-17]⭐⭐ PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the corpus is clean. Operator: "he's a great writer because of his research, not because he has a strong identifiable voice." I surveyed, built the corpus, measured containment and fixed three stripping defects before anyone asked the question that decided it. henge id 82. → auto-memoryfeedback_voice_worth_adapting_before_corpus. -
[2026-09-17]The althing route-declaring SessionStart hook is documented but NOT installed on nh3-dev —dev_launch.pyhas zero occurrences of "route", no hook declares one, and every live route was hand-declared at a different minute. A relaunched pane therefore drops to pull mode and stops receiving mail. Raised with the operator, not acted on, because installing it edits~/.claude/settings.jsonand a peer's question is the wrong authorisation for a config change. Tracked at althing thread01M2R0KPPE85SVJ96KSQ3YKKQP. -
[2026-09-17]Hemingway ships as-is: operator ruled "ship stands" on both measured corpus defects — the 0.96% beat contamination and the 130 non-name entity-map surfaces.audit_pairs_sourcenames.py --filter-outandaudit_entity_map.pyexist and are the instruments if that is ever revisited; neither was run against the shipped adapter. -
[2026-09-17]⭐⭐⭐ PARKED lv-krakauer, and the reason is a selection criterion the line was missing: ask whether the author HAS a voice before investigating whether the corpus is clean. Operator: "he's a great writer because of his research, not because he has a strong identifiable voice." A voice adapter is worth its ~6 hours only if the target has a prose signature a reader could pick out blind; Krakauer's excellence is reporting, which an adapter cannot carry. I surveyed, built, measured containment and fixed three stripping defects before anyone asked the question that decided it. For each candidate, say what the voice IS in one sentence and how it shows up in char-bigram space, before the first catalogue query. henge id 82. → auto-memoryfeedback_voice_worth_adapting_before_corpus. -
[2026-09-17]⭐⭐ A unit splitter must choose by SIZE, not by count — the inherited rule silently produced 22,000-word "chapters".scripts/r49-corpus/split_units.py: a marker mode qualifies only if its median unit is inside [600, 12000] AND no unit holds half the work; among qualifying modes PRIORITY breaks the tie (contents > chapter-word > roman > bare-numeral > caps-title), and paragraph-block sections are the fallback for works with no divisions. ⭐ Both rules exist because a control caught them: scoring by "median closest to target" chosecaps-title(6 units, one holding 97% of the book) over True at First Light's real 20 chapters, because a median cannot see that distribution and a max bound can. Positive control: 8/10 Hemingway works reproduce the shipped mode and count exactly. Negative control: 40,000 words with no blank lines → 1 unit, refuses to fabricate divisions. Commit705fa3a. -
[2026-09-17]⭐ lv-mccarthy D1 built — 167 units, 584,756 words — and the whole job was protecting a style that reads as damage. 0.0 quote marks per 10k (Hemingway 838),dont/aint/wont. The builder runs NO typography normalisation and asserts the quote density afterwards. Two truncated catalogue rows dropped for complete mobi siblings; all 15 containment pairs measured (worst 0.10%); back matter in 4 of 6 works carried the author's name 26 times → 0. ⚠ The back-matter strip runs BEFORE the split here — Blood Meridian and The Crossing end with a dumped TOC of bare roman numerals, the exact shape of a chapter marker. Commitf3bf3ca. -
[2026-09-17]⭐ lv-krakauer D1 built — 126 units, 422,880 words — and its name guard caught three defects nothing else would have reported. Back matter searched only the LAST unit while the apparatus sat in unit 37 of 41; relying on the splitter to drop front matter failed because the ebook TOC sits above the author's note and gave it aChapter Thirty-Twoto start on; and zero was the wrong bar — 2 survivors are Krakauer writing about his own father in Into the Wild's autobiographical chapters, so the allowance is pinned at 2 with every survivor printed. ⚠ Both strips are windowed in the OPPOSITE direction from McCarthy's, because Krakauer'sALSO BY/Copyright/About the Authorsit at 0.0–0.6% of the file. Commit4be0630. -
[2026-09-17]⚠triage_disposition = 'accepted'in the Kvasir catalogue does NOT mean the extraction succeeded. Blood Meridian's epub row holds 1,167 words of a 117,000-word book, The Crossing's 222 of 150,000 — real prose, real titles, accepted. Faulkner's The Mansion is 39 words.near_dup_pairsholds ONE row in the entire 1,284-work library and is blind to a fragment beside its full sibling. Word-count every master before trusting a row, and note that word count alone cannot tell a truncated novel from a legitimately short work. -
[2026-09-17]⭐⭐⭐ lv-hemingway SHIPPED (ckpt850) with the line's strongest voice result — and the memorisation control it passed turned out to be the WRONG control. Voice +0.413 delta_cb at 6.4x the floor, closing 73.8% of the achievable span (lv-bronte closed 48%). ⚠memorization_check.pyuses the base-unadapted arm as its negative control, but base writes 18,035 words of summary against the adapted arms' 27,413 of pastiche — text that does not imitate the register cannot collide with its n-grams, so a 0.00 there means "different register", not "did not memorise". The right reference is the author himself: held-out Hemingway against the train split collides at 0.01 while the adapter does at 0.07, so the comfortable "his plain register makes collisions inevitable" story is FALSE and was refuted rather than assumed. All 19 matched runs were READ: stock dialogue, max 9 words, no proper noun — shorter than the 10-word run unseen Hemingway shares with the train split by coincidence. ⭐ A negative control that differs from the candidate in a way correlated with the metric is not a control. →persistent-memory.d/2026-09-17-lv-hemingway-gate.md -
[2026-09-17]⭐⭐ The v2 voice floor is now PAIRWISE, and it retroactively passes lv-bronte. lv-bronte's ckpt475 shipped as a voice-axis FAILURE at +0.193 against a 0.251 floor contributed entirely by ckpt925 — a third arm nobody was shipping, on one outlier seed. Scored against the arm it was actually compared to the floor is 0.091 and it clears at 2.1x. The rule was changed prospectively, pre-registered for lv-hemingway before any Hemingway number existed, on an argument independent of the answer: the sampling variability of a difference A−B depends on A and B, never on a third arm C. The previous session found the defect and deliberately declined to exploit it; this follows from fixing it. lv-hemingway passes under both rules, so its verdict does not lean on the change. Caveats amended append-only in the compose, the NFS README and the gate record. Commits0bb49382e9b118. -
[2026-09-17]⭐ The beat-contamination leak IS present in Hemingway — 70 of 7,094 train beats (0.96%), 0 of 200 val.scripts/r49-corpus/audit_pairs_sourcenames.pycloses the blind spotleak_gate.pyhas by construction (it reads the corpus and the renamed copies, never the generated beats). Cross-validated on real data: the fixed Brontë pairs return 0 of 3,858 andpairs-full.CONTAMINATEDreturns 15 of 792 = 1.89% with the recorded names.--filter-outyields a verified-clean 7,024-pair set in one command; the retrain is the operator's call. The val split being clean is why the gate could run at all. -
[2026-09-17]⭐audit_entity_map.py— the rename can DAMAGE the prose and no gate will ever say so. Mirror ofaudit_stoplist.py: surfaces wrongly held IN the map rather than out of it. 130 of 946 Hemingway surfaces flagged (1,616 instances, 0.162% of words) —African,Chinese,X-ray,Coca-Cola,Ritz,Pradorenamed into invented names — plus 16 bare initials incl.Cat 274 occurrences. Signal is a preceding article; controls derived from the corpus, not hand-picked. Every hit reported for READING:the Widowandthe Informerare genuine epithet-names that should be renamed. Commit051b99e. -
[2026-09-17]The two-epoch recipe is now 0 for 2 and should stop being carried forward. Hemingway's eval minimum is step 1750, but step 850 is +0.0040 against a 0.0044 median neighbour jitter — three checkpoints inside one jitter — and 850 won every resolving axis (2.3x tighter seed spread, lower memorisation, less ran-on). Same outcome as Brontë. What IS robust on this schedule is the epoch-3 collapse: +0.0762 = 17.4x jitter. -
[2026-09-17]gitea was reaching the PUBLIC route from every repo on nh3-dev — brokkr-smithy, sleipnir, Galdrabok, kvasir — and brokkr-smithy is pushed several times a week, so the fail2ban trigger was live, not dormant. Measured before acting (no split-horizon rewrite, no ssh alias,ssh -Gconfirmed port 22 to 38.120.12.44). Fixed by overriding the NAME once in~/.ssh/configrather than rewriting N remotes, so fresh clones and unaudited repos are covered too. Verified with a realgit ls-remote, not by inspection. Commitdcc1abc. Flagged by brokkr-smithy-dev;vh/imogencreated for them the same session. -
[2026-09-17]servers/fv-ml1/ssh-targetwas bare10.251.50.54, sodeploy-stack.shconnected aslkravenand could not write the infra-ops-owned/opt/docker/compose/— and lkraven's sudo on fv-ml1 needs a password, soDEPLOY_SUDO=1failed too. Nowinfra-ops@10.251.50.54;--validate-onlystill clean, deploy works. ⚠ Other hosts'ssh-targetfiles may carry the same gap — a read-only refresh works as either user, so the fault only surfaces on a deploy. -
[2026-09-17]⭐⭐ A leak class the corpus gate structurally CANNOT see: the beat-writing model recognises the book and restores the author's real character names. 1.8% of Brontë beats named Rochester/Jane/Brocklehurst while 0 responses did. Worst for public-domain classics; Yarros and Hemingway's clean runs are NOT evidence they are immune. Patched as asourcenamereject +--source-entities. →persistent-memory.d/2026-09-17-beat-contamination-leak.md -
[2026-09-17]⭐ A stoplist entry is an assertion the leak gate can no longer check — stoplisting removes a surface from the entity map, so a wrongly stoplisted CHARACTER is an undetectable leak. Three were wrong on Brontë (Leaven, Pierrot, Samuel);scripts/r49-corpus/audit_stoplist.pyfinds them by honorific and now gates the pipeline. Commit8bb7686. -
[2026-09-17]ESH: Cityside Fiber failed TWICE (19:09 and ~01:06); operator switched WAN1 to DHCP to restore service and has a ticket for the static. crowdseceshallowlist carries both failover egresses with 7-day expiries — the rotation-fragility is live. →persistent-memory.d/2026-09-17-esh-fiber-outages.md -
[2026-09-17]⭐⭐ lv-bronte SHIPPED on voices-seat (ckpt475) DESPITE failing the v2 VOICE axis — additive, reversible, safety-axis clean. Both candidates closed 48–52% of the achievable distance to Brontë but +0.193/+0.210 sit under a 0.251 noise floor set by ONE outlier seed in the arm not being shipped; cause is structural (81 val pairs vs Hemingway's 200) and not cheaply fixable. ckpt475 is the pick if it ships. The two-epoch recipe did NOT transfer. →persistent-memory.d/2026-09-17-lv-bronte-gate.md -
[2026-09-16]⭐⭐ Instruction-pair SFT BEATS raw-text for author voice, and the raw-text incumbent never cleared its own control (+0.141 against a 0.153 floor).lv-yarrosshipped; the v1 decision rule was amended by the operator after measurement showed it gated on axes the unadapted carrier already maxes. →persistent-memory.d/2026-09-16-lv-voices-line.md -
[2026-09-16]⭐ voices-seat live: one carrier, Nlv-<author>LoRA adapters, hot-swap measured at 0.24 s. LoRA costs 24.3% of decode against a 0.1% A-vs-A floor and is worth paying;--gpu-memory-utilizationis a request against TOTAL VRAM and only a pinned KV makes it predictive. →persistent-memory.d/2026-09-16-voices-seat-lora.md -
[2026-09-16]⭐ lv-hemingway corpus gated at 994,760 words — and half the catalogue had to be EXCLUDED. 169,759 words of measured 90–96% collection duplication, a Sherwood Anderson parody, and the author's own name 95 times in publisher back matter; the gender resolver needed a corpus base-rate correction to stop reading women as men. →persistent-memory.d/2026-09-16-lv-hemingway-corpus.md -
[2026-09-16]Grok token broker built then SHELVED — operator ruled "keep the jail stop the a/b", so the renewal feature has no consumer. ⛔ Do NOT armprobe-rotation: the risk did not shrink (it reaches BOTH Gróa transports through one shared session) and the payoff went to zero. →persistent-memory.d/2026-09-16-grok-broker-shelved.md -
[2026-09-15]⚠⚠ DO NOT carry "a client-side timeout is not a cancellation" as a rule — it is FALSE as stated. A clean abandon cancels itself ~6 s later (measured); yet six requests genuinely orphaned onvllm-erp-seat. Some propagate, some do not, boundary unknown — which argues for a detector, not a rule. ⭐⭐ The durable artifact: a serving engine's KV cache CYCLES, an orphaned one only CLIMBS — request count and throughput are ambiguous between loaded and wedged, and I called the seat healthy twice off them (correctly, on the evidence). ⚠ Amax_tokensceiling would NOT have prevented it: the worst offender had 16384 set, hit it, and returned 24,594 chars of whitespace. →persistent-memory.d/2026-09-15-client-abandon-cancellation-boundary.md -
[2026-09-15]⚠⚠--gpu-memory-utilizationDOES NOT PREDICT RESIDENT VRAM — measure it, never compute it. Wrong in both directions on fv-ml1:vllm-cyberprevutil 0.40 (expect ~39,155 MiB) holds 47,124 (+8 GB over);vllm-gen-smallutil 0.48 (expect ~46,986) holds 36,942 (−10 GB under). Planning a placement off the fractions would have been 8 GB wrong. Readnvidia-smi --query-compute-apps. Full per-seat residency table + the breeze shuffle arithmetic →persistent-memory.d/2026-09-15-breeze-placement-sizing.md -
[2026-09-15]breeze-tts stays on irv-ml1; the TTS-stack move to fv-ml1 is PARKED (park id 75,move-the-tts-stack-breeze-tts-bragi-tts-gateway), triggered on evacuating embed/rerank/reward. ⚠ Trigger as stated says "gpu0" but those three are on GPU 1 (~0.16 util, ~15.7 GB; GPU 1 is the tight card at 0.975 / 4,336 MiB free) — confirm which he meant before executing. All three services move together because onlybreeze-ttsis GPU-resident (~10.3 GiB, growing) whilebragiandtts-gatewayare CPU proxies, and co-location is what avoids a cross-site hop per TTS call. breeze-tts sizing — original recommendation NOT to move it. ~10.3 GiB measured under load at 53 min uptime, up from 9.2 GiB shortly after warm-up (it grows; n=2, plateau unmeasured) — so GPU 0's 11,982 MiB free is a 1.7 GB margin and shrinking, on the live chat serving path. ⚠ Two measurement traps: it reports nothing at idle on the wrong card (BREEZE_GPU_DEVICES=0= the 3090, not the A6000), and an early reading understates it. ⭐ The real objection is topology:tts-gatewayis on irv-ml1 and reaches it same-box, so moving breeze alone adds a cross-site hop to every TTS call against a 478 ms first-sample budget. GPU 3 would fit it but spends the reserve. →persistent-memory.d/2026-09-15-breeze-placement-sizing.md -
[2026-09-15]Parakeet STT live on fv-ml1 GPU 0, behind LiteLLMext-stt/whisper-1. ⚠ Placed on GPU 3 first, which was wrong — operator caught it. A ~800 MiB seat should ride the card with the most uncommitted headroom (GPU 0, util 0.88, ~13 GB spare), not put the first fingerprint on the one pristine 96 GB card: vLLM sizes KV cache against TOTAL VRAM, so any tenant on an empty card eats a future full-size seat's profiling margin (flash-next needs 93 of 96 GiB). GPU 3 is now a deliberate reserve at 2 MiB. Retargeted the existingstacks/parakeet/(sherpa-onnx + our own FastAPI wrapper) from irv-ml1; v3 int8, 25 languages. ⚠ ORT's CUDA EP compiles kernels lazily and the first decode on sm_120 took 45.7 s — every later call ~0.5 s; a startup warmup inapp.pynow absorbs it, so the first real request is 0.65 s instead of a 45 s hang that no client would wait through. GPU use was verified by a process on GPU 3 (922 MiB), not by theprovider=cudalog line, because ORT falls back to CPU silently and still returns correct text. Silence →""(null control), known sentence → near-exact (positive control). →persistent-memory.d/2026-09-15-parakeet-stt-fv-ml1.md -
[2026-09-15]⭐⭐⭐ THE FLEET'S CHARACTERISTIC FAILURE, named: a confident answer from a broken instrument. Nine instances in one night, every one of which PASSED A CHECK —provider=cudawhile ORT ran on CPU;node --checkgreen on a file whose SERVED script was dead;secret getreturning""with exit 0;find()turning a failed listing into an authoritative "not found"; a 401 rendering as "0 toolsets";compat✓ on a typo'd path;doctorexit 0 on ERROR;ss | grep pythonmissing a listener namedhermes; SIGTERM freeing a port 35 s before the process died. ⚠ The tell: whenever "broken" and "legitimately empty/absent/off" produce the same output. Remedies: measure the output not the input, positive AND true-negative controls, refuse to emit the ambiguous value, and never declare victory on a plausible fix. →persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md -
[2026-09-15]secret getreturned EMPTY with exit 0 under concurrency (svos-dev found it; 0/4 succeeded here). Root cause isbw unlockracing at session establishment, not item reads — so a lock inside the read wrapper cannot work. Fixed: command-level lock,cmd_getrefuses an empty value, andfind()no longer coerces empty stdout to[]. ⚠~/.local/bin/secretwas a plain COPY — now a symlink.0193b31. -
[2026-09-15]⭐⭐ A check that reads an artifact AS STORED cannot see a transformation between storage and execution — named twice in one night and it generalises.node --checkon a source file passes while the SERVED page's inline script is dead (a JS'didn\'t'inside a Python string arrives as'didn't'and closes it);provider=cudain a log echoes configured intent while ORT silently ran on CPU. Both check the INPUT to a transformation and get reported as checks of its OUTPUT. Remedy: gate the wire, not the file —tts-stack tools/gate_served_page.py. ⚠ My first version had a gap tts-dev closed: a worklet inside a template literal is just a string to a parse of the enclosing script, so its syntax error surfaces as a rejectedaddModulepromise and silent degradation. I checked the instance, not the class. →persistent-memory.d/2026-09-15-talk-v10-deploy.md -
[2026-09-15]⚠⚠ The talk-deploy "permission problem" NEVER EXISTED — and I built a fix for it anyway./opt/docker/composeon nh3-dev isroot:docker 2775, sessions run aslkraven,lkravenis indocker; amkdirsettles it in one second and nobody ran one for nine days. There is notts-devOS account at all. It held because a stale memory row supplied a mechanism, the operator's routing instruction ("give it to infra") was misread as corroboration of a capability limit — different claims, only one ever stated — and I repeated it to the operator as fact. Then, told to fix "the harness issue", I inferred an auto-mode classifier refusal and committed a settings.json to tts-dev's repo on that inference; theirmkdirdisproved it and I reverted. ⭐ "I can't do X" is a hypothesis until someone pastes the error. ⚠ That commit also overclaimed a doc fix that failed — never chain an edit and its commit in one invocation. →persistent-memory.d/2026-09-15-silent-wrong-answer-pattern.md -
[2026-09-15]talk v10 LIVE on nh3-dev :8092 — the fleet speaks and listens on one page. First consumer of theext-sttParakeet seat:POST /api/listen, push-to-talk, barge-in. Gated build→throwaway→teardown→cutover, then re-gated against production (a gate that only ran against the throwaway proves the image, not the deployment). ⚠ Deploys route through infra-ops only because tts-dev's identity is not in nh3-dev'sdockergroup — a permissions accident, not a judgement call; group-vs-relay is in front of the operator. -
[2026-09-15]⭐⭐ Two restart patterns from svos-dev worth stealing: (a) DRY-RUN BOOT against the still-held port — start the new process while the old one holds the socket; it proves every check above the bind and dies on[Errno 98], so a one-way restart becomes a rehearsed one at zero cost. (b) ⚠ SIGTERM freed the port but left the process alive for 35 s — a script waiting on the port would have run two copies. Kill by PID, wait on the PID, never on the port. A freed port is not evidence of a dead process. -
[2026-09-15]⭐svos_mirandaENABLED and LIVE in Hermes — butagent.disabled_toolsetsis permanently OFF by operator ruling ("i dont want the tools disabled everywhere"). That key is a global end-of-pipeline subtraction, not api_server-scoped: measured 46 tools → 20 on a default session. It is also unnecessary —platform_toolsets.api_server: [svos_miranda]alone resolves an api_server session to exactly the 8 tools, write-klass absent. Gateway restarted 02:10 (PID 3107822→3901622, observed);/v1/toolsetsnow 29 rows incl.svos_miranda; operator's own surface verified intact at 46. ⚠ SVOS must stop verifying against the GLOBAL roster before it restarts — it will see 29 and refuse, by design now. →persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md -
[2026-09-15]irv-ml1 parakeet RETIRED; voice-studio STOPPED. Both operator rulings. Parakeet lost tts-dev's bench to the FV seat at both clip lengths and to whisper-large-v3 at 6.24 s; no gateway alias depended on it and every other host reference was a port-register comment. voice-studio existed for the dots mint loop, which Breeze obsoleted 2026-09-06 — retired rather than repaired. -
[2026-09-15]svos_mirandaHermes plugin validated; found its load blocker. Absolute intra-package imports (from hermes_plugin.x) could not resolve at the documented install name — fixed by svos-dev atc964e64. ⚠hermes plugins validateanddoctorcan NEVER pass this plugin, by construction: validate's probe stub is config-blind AND returnsNonefromregister_tool(which the plugin's guard reads as a collision), and doctor runs under a tempHERMES_HOMEwith no config. ⚠doctorexits 0 on ERROR (use--ci);compatreads a nonexistent path as a pass. Roster verified 8/7 by a probe supplying real settings. →persistent-memory.d/2026-09-15-svos-miranda-plugin-validation.md -
[2026-09-15]⚠ ana-docker resolves NO.internalnames — its/etc/resolv.confis1.1.1.1/1.0.0.1, not the fleet AdGuard. LiteLLM only reachesirv-ml1.nh3.internalbecause of a hand-pinnedextra_hostsin its compose. New gateway aliases therefore use raw IPs; adding a hosts entry would mean recreating the container and bouncing the gateway for every consumer. Fleet-wide DNS fix is unowned. -
[2026-09-15]⚠⚠ irv-ml1 still points at the retired wg0 lifeline10.100.79.3in 96 places — and one is a LIVE breakage, not a dead link.voice-studiocannot reachstudio-gate(both up, separate docker networks, gate URL is the dead IP) and has been failing since the 2026-09-06 cutover with nothing alerting. 8 running containers carry deadhomepage.hreflabels;waterland-studio's siteMonitor too. ✅tts-gateway/ext-ttsverified UNAFFECTED. Not fixed — wants a scheduled pass, not a 02:00 improvisation. ⭐ Third instance of the same shape: a retired address needs a grep by ADDRESS, not by hostname, and labels live in no file until the container is recreated. →persistent-memory.d/2026-09-15-irv-ml1-dead-wg0-address.md -
[2026-09-15]Parakeet bench settled by tts-dev — FV wins at both clip lengths and beats the incumbent Whisper; IRV seat is now retirable. FV 155 ms / 391 ms on 1.84 s / 6.24 s clips vs IRV 354 / 1010 vs whisper-large-v3 457 / 690 — IRV is slower than Whisper at 6.24 s. Length sweep (n=9/cell, first 3 discarded) fits ~58 ms fixed + 56 ms per audio-second, asymptote ~17.8x realtime, which independently reproduces our 17x on a different harness. Gateway hop measured below harness resolution (±30 ms), soext-sttis the right consumer path. ⚠ tts-dev retracted their own plan's 60-120 ms projection: published RTFx is BATCHED THROUGHPUT, not single-stream latency — the two differ by ~200x. ⚠ Their between-run variance is ±20% because GPU 0 is the live chat path; our 0.50 s median was taken on an idle GPU 3 and is a best case. -
[2026-09-15]Mesh membership retired for fv-ml1 and nh3-dev — six nodes left, each with a job. fv-ml1 gets break-glass rejoin instead of standing membership; nh3-dev's retirement also removed the nh3-scale masquerade exception it had required. Exactly one live reusable pre-auth key remains fleet-wide. →persistent-memory.d/2026-09-15-fv-mesh-watchdog.md -
[2026-09-15]FV cross-site routing fixed — one OPNsense outbound-NAT rule had been scoped to Anaheim only. fv-ml1 now reaches NH3/ESH/IRV/ANA/mesh/internet; four rules, allsrc=10.251.50.0/24. The diagnostic signature is the valuable part: every layer looks correct and the discriminator is that every other site pair works. →persistent-memory.d/2026-09-15-fv-cross-site-snat.md -
[2026-09-15]Break-glass mesh path on fv-ml1 — inverted from a restore-watchdog on the operator's suggestion: the box is OFF the mesh and the watchdog JOINS it on fleet loss. Exposed a rejoin key expiring in 4 days; replaced with a dedicated 1-year key and the two stale reusable keys retired. →persistent-memory.d/2026-09-15-fv-mesh-watchdog.md -
[2026-09-15]Fleet identity/group/path conventions pinned + docker trees →root:docker 2775setgid on 5 hosts.svc-*in 800-849, infra-ops 850, docker 851,vhfor new hosts with no retro-renames;0777cleared;linusdeleted;llmuserde-privileged. →persistent-memory.d/2026-09-15-fleet-identity-conventions.md -
[2026-09-15]nh3-dev unreachable from the mesh at its LAN address — Tailscale'sts-inputanti-spoof, not DNS. Fixed with a masquerade exception on nh3-scale. ⚠ Do NOT instead advertise the /32 from nh3-dev; that black-holes it from every other site while its own LAN keeps working. →persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md -
[2026-09-15]ESPHome pinned to 2026.8.2 +kbKB-search tool shipped. Untagged image had drifted a year; config relocated into restic with 539 MB of regenerable cache excluded; remote-build disabled (⚠ two switches, only one closes the port).kbexists because Worldtree's/searchsearches messages, not notes, and returns a clean empty result for a note that exists. →persistent-memory.d/2026-09-15-esphome-and-kb.md -
[2026-09-15]Hermes bearer rotation hold released — svos-dev split their HS256 signing key off the shared value (svos7165272). ⚠ When thesvos_mirandaplugin arrives it will reference the dispatch key, not the bearer (expected), and itstoolsarray is legitimately seven or eight entries; any other number is a real fault. Commite641931. -
[2026-09-14]fv-ml1 rebalance: cyberprev→sec(mog-sec retired), NEW gen-small A3B seat, all sec/gen/char at native 262K in-band, coder reclaimed, seat catalog + bench shipped. cyberprev = hotdogs cyber-SFT (name-repaired past a tripled-prefix unsloth export bug, house NVFP4 quant); gen-small = llmfan46 Qwen3.6-35B-A3B Heretic (already on disk), MTP 69.6%. Serial depth-tested all seats clean (0 OOM); warm tok/s 62.7-337.3. Commits 1418edb→dfa91a8. →persistent-memory.d/2026-09-14-fv-seat-rebalance-gen-small.md -
[2026-09-14]fv-ml1 all-night seat reorg — MTP k=3 on gen-large (+52%@conc1), gen consolidated onto flash-next (27B dense retired, 38 GB freed), char-rp restored to MeroMero-v2-31B, Sentinel-R3 served + dflash cutover (beat MTP 2.40 vs 2.18). ✅ gen-large RESOLVED 2026-09-14 — orcarouter serving: PLE bf16→FP8 convert +ple_embedding_dtype+layer_typesrename; NO source build needed. →persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md -
[2026-09-13]STANDING POLICY (operator): cap GPU power limits at BUILD time, not after discovering the constraint. →persistent-memory.d/2026-09-13-standing-policy-operator-cap-gpu-power.md -
[2026-09-13]**FV SITE DARK — every Fountain Valley address including the BMC went unreachable ~2.5 min into a two-card ** →persistent-memory.d/2026-09-13-fv-site-dark-every-fountain-valley.md -
[2026-09-13]⭐⭐ Qwen3.8-Flash-Next serving on ONE card with its 51B n-gram table in host RAM — the first seat whose weights do not fit its GPU.stacks/flash-next-seat/, fv-ml1 GPU 2:8022, plus agen-largeLiteLLM alias. Measured: 74.36 GiB weights resident, 14.00 GiB KV = 560,654 tokens at the full 262,144 context, 67 GiB host RSS, 75.5/212.3/387.8 tok/s at conc 1/4/8 (⚠ n=1). ⭐ The offload is vLLM #54371 (UVA, merged 2026-09-09) which supersedes the paused #53899 — it has no worker process, so #53899's whole bug family (TP=1 deadlock #53960,pidfd_getfd/ptrace gate, stale-output-under-graphs) is designed out; inv0.29.1rc0, notv0.29.0. ⚠text_config.ple_embedding_dtypeis the load-or-fail discriminator for any community build. ⚠⚠--kv-cache-memorymakes vLLM SKIP MEMORY PROFILING and ignore--gpu-memory-utilization— 16 GiB nearly OOM'd on a 155K prefill with no visible failure; 14 GiB is the measured-safe value and vLLM's own "17.46 GiB to fully utilize" is 3.5 GiB too high. ⚠ MTP is off pending measurement here, not written off — the recipe's number is cross-harness and tested k=3 only, while the head is ONE layer run autoregressively, so k=1 is unpublished and may win (services/flash-next-mtp-bench/, oneoff_Arep banked before the outage). ⚠ A container once ran(healthy)withPORTS=[]— verifydocker port, not the healthcheck. →persistent-memory.d/2026-09-13-flash-next-seat-and-fv-outage.md -
[2026-09-13]Finished the ana-ml2→fv-ml1 renumber the cutover missed: 16 live Homepage entries pointed at the dead 10.250.50.54 and zero at the live IP. The sweep allowlist was built from files that mention the HOST and ahomepage.hrefmentions only an IP, so every label-only stack fell outside it by construction; 24 files fixed plus host copies, allowlist extended with how to derive it next time. Two bugs fell out:deploy-stack.shrejected any stack name containing a dot (soqwen3.5-122b/qwopus3.5-122b/mistral-medium-3.5could not be deployed at all), and scriberr's CORS allowlist held only the dead IP and the deadscriberr.ana.internal. ⚠ Incomplete: the 10 running containers were never recreated and a plain power-on will NOT apply labels — the stagedcompose up -drecovery does. Eight stacks deliberately not pushed (real host drift); three of those are untracked host-only stacks. Commits3132a16,969a1b6,d79f104. -
[2026-09-13]FV→ANA fixed, Beszel18/18 up: scoped OPNsense hybrid NAT for fv-ml1→ANA; prior NAT/filter rules preserved and rollback-guarded. →persistent-memory.d/2026-09-13-fv-to-ana-nat.md -
[2026-09-12]⭐⭐⭐ FV CUTOVER EXECUTED — fv-ml1 live at Fountain Valley, renamed, renumbered to 10.251/16, serving inference; BMC online after finding it was tagging 802.1q VLAN 250 into an untagged port; and the box has FOUR RTX PRO 6000 (391 GB VRAM), not the two every doc claimed. Also: OPNsense write APIs need anX-CSRFTokenheader scraped from a<script>block — session cookies alone 403, which is why reboots appeared to work while the operator was power-cycling by hand. →persistent-memory.d/2026-09-12-fv-cutover-executed.md -
[2026-09-12]esh-vm-db Restic fixed: stale April PG dump caused by TCP auth + masked errors; peer-auth/fail-closed hook and bounded retries installed. Snapshot bc5eeaff and repository check verified. →persistent-memory.d/2026-09-12-esh-vm-db-restic-repair.md -
[2026-09-11]Worldtree memory-split (U6) — PROTOCOL AGREED with worldtree-dev: nobody flips memory.reader.enabled or m →persistent-memory.d/2026-09-11-worldtree-memory-split-u6-protocol-agreed.md -
[2026-09-11]⭐ Plex hardware transcoding on the Arc A580 FIXED (esh-pve-nas LXC 105) — every setting was already correct and the fault was one layer below them.intel-media-va-driver22.3.1 (Apr 2023, stock jammy) predates Arc/DG2 support and exports only__vaDriverInit_1_14, against the libva 2.22 Plex BUNDLES and loads via RPATH. Passthrough, cgroups,plexin video+render, HuC authenticated, Plex Pass,HardwareAcceleratedCodecs=1and the Arc already selected asHardwareDevicePath— all good the whole time. Fixed with Intel's client-GPU repo (rollingjammy client) → iHD 24.3.4 (__vaDriverInit_1_22) + a consistent libva 2.22.0.2-87 set, now pinned +apt-mark hold(verified: a simulated upgrade moves 152 packages, touches none of the six). Also repaired a half-finished prior attempt — libva/libva-drm hand-installed at 2.22 withlibva-x11left at 2.14, killing every X11 VA-API app onva_fool_postp. ⚠⚠pct snapshotREFUSES on a bind-mounted guest AND STILL EXITS 0 (LXC 105 hasmp0: /tank/media) — usezfs snapshot nvme/subvol-105-disk-0@<tag>and read it back. ⚠⚠ A syntheticPlex Transcoderrun is NOT a valid test (Plex bundles its own libc among 61 libs; my harness failed identically before and after a fix that worked — no positive control, so its negatives were worthless). Only a forced transcode settles it: PASS names the device (testing API vaapi for device '/dev/dri/renderD129' (Intel DG2 [Arc A580])). ⚠ The original emptyfinal decoder: , final encoder:was an absence of evidence, not failure —TranscodeSessionwas 0. Jellyfin LXC 107 left alone (operator: not actively used). →persistent-memory.d/2026-09-11-plex-arc-vaapi.md, runbookdocs/runbooks/plex-arc-vaapi-jammy.md -
[2026-09-11]Beszel priority 2 complete: both DB hosts and both PBS hosts verified, 16 new alerts; fleet 17/18 up (ana-ml2 down). →persistent-memory.d/2026-09-11-beszel-priority2.md -
[2026-09-11]Beszel priority 1 complete: all six installed and verified. After Anaheim recovery, live Synology samples and alerts verified; fleet 13/14 up, only known ana-ml2 outage remains. Configs not committed. →persistent-memory.d/2026-09-11-beszel-priority1.md -
[2026-09-11]Sentinel-R3 pulled, MTP-grafted, and quantized as a M.O.G.-SEC seat candidate — quant DONE, acceptance UN →persistent-memory.d/2026-09-11-sentinel-r3-pulled-mtp-grafted-and.md -
[2026-09-11]MEASURED: two concurrent training jobs on pfi-gx10 are 13% NET SLOWER than running them back to back — VR →persistent-memory.d/2026-09-11-measured-two-concurrent-training-jobs-on.md -
[2026-09-11]ana-ml2 → fv-ml1: relocating to a NEW Fountain Valley colo TOMORROW (operator decision). Its power draw ( →persistent-memory.d/2026-09-11-ana-ml2-fv-ml1-relocating-to.md -
[2026-09-11]Anaheim rack LEFT DARK until the move (operator decision). ana-ml2 is the ONLY host still down post-recovery (BMC dark = no power); rather than power it on tonight just to shut it down for the truck tomorrow, it stays off. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) but there is nothing to bring up — the box relocates as fv-ml1. -
[2026-09-11]RECOVERY FOOT-GUN, will recur every colo power event: crowdsec crashes on the hard power-off and traefik' →persistent-memory.d/2026-09-11-recovery-foot-gun-will-recur-every.md -
[2026-09-11]Anaheim colo recovered ~16:39 PT EXCEPT ana-ml2 (bare metal, NO power — its BMC 10.250.250.50 is dark on standby, unlike same-subnet pfi-pve which is up → needs a physical PDU/PSU/breaker fix, not a boot). pfi-pve + all its VMs (ana-docker/ana-nas/ana-wg/corviduo-dev/pbs-ana) auto-started clean (on-boot gap held this time). LiteLLM came back up on its own (transientunhealthyduring startup → serving). ⚠ Public WAN (38.120.12.44) ICMP still blocked from outside but HTTPS works fleet-internally (mesh-routed). ana-ml2 down blocks the gen/summarizer/mog-sec seats AND the cyber-preview quant re-run. I hold vaulted IPMI creds (ana-ml2/bmc-{infra-ops,password}) to power-on + boot-watch the instant its BMC returns. -
[2026-09-11]BabyYarros COMPLETE — both arms trained AND evaluated; the voice moved toward Yarros above the measured n →persistent-memory.d/2026-09-11-babyyarros-complete-both-arms-trained-and.md -
[2026-09-11]⭐⭐ BabyYarros UNBLOCKED and TRAINING: the leak gate passes at 0 of 325 entities and 0 of 91 phrases, and closing it turned up three defects nobody was looking for. The gate itself is the first artifact — there was no committed instrument for "does any of the author's proper nouns survive", so Brontë's 0-of-203 was a hand count.scripts/r49-corpus/leak_gate.pynow runs the same scan over the UNRENAMED source as a positive control plus a nonce negative control every time, because a detector that only ever sees renamed text cannot tell absent from blind. Its first reading was 212 surviving, not 86 — it scans the whole corpus rather than per work, and counts the sub-threshold entities rename never looked at. Training launched 10:06 PT on pfi-gx10: Qwen3-4B-Instruct, 1 epoch, seed 4919, 178 steps / 5,824,512 tokens, corpus shae85f69f1e49d57c9. →persistent-memory.d/2026-09-11-babyyarros-leak-gate-passes.md -
[2026-09-11]**A SECOND corpus typography defect, and the D1 note that "no unwrap was needed" was right about the wrong ** →persistent-memory.d/2026-09-11-a-second-corpus-typography-defect-and.md -
[2026-09-11]⚠⚠ Back matter was inside the prose of all five works — 4,555 words naming the author's agent, editors and children. The builder splits on chapter headings and nothing follows the last one, so acknowledgments, newsletter pitches and cover-artist credits rode inside the final chapter. Found by the gate's phrase audit surfacingLouise Fury(Yarros's literary agent), not by reading. ⚠ iron-flame's marker isACKNOWLEDGMENTSin all caps and a case-sensitive scan missed it — the strip is case-insensitive and last-chapter-only, with an acceptance check that refuses if it would remove more than 2% of the corpus. -
[2026-09-11]⭐⭐⭐ The gate read 0 of 314 whileAfendrawas still in every copy — the worst failure shape available. The name never appears unpossessed, so it keyed asAfendra's, and rename.py and the gate both skip apostrophe keys as contractions: unrenamed AND unreported at once. Fixed by folding clitics (--fold-clitics) soAfendra'scounts asAfendra.Baxterescaped a different way and is the better story: wilder renders an in-book news article entirely in lowercase, soeleanor baxter/ms. baxterappear uncapitalised 3 times against 23 capitalised — ratio 0.13 against a 0.05 bar, and a real character is silently never renamed. Fixed by readmitting ratio-rejects that a title precedes (--rescue-honorific 2). ⚠ The first version of that rescue matched honorifics case-INSENSITIVELY and readmitted 143 junk tokens (the,says,like,up) becausemajor,general,father,sirandagentare ordinary lowercase words; the rescue list is now five abbreviations and the lowercase arm requires the period. -
[2026-09-11]⭐⭐ A whole leak class the unigram scan structurally CANNOT see:Riders Quadrant,Flame Section,War Games— andFourth Wing, the book's own title. Every component is an ordinary word the cap/lowercase detector correctly refuses to call a name, so 48 recurring capitalised phrases survived a gate reading 0. This isThornfield × 100one level up, and it needs a map, not a detector — substituting a head noun is a choice about register, not a measurement.scripts/yarros-corpus/phrase_map_yarros.json(10 phrases + 13 capitalised tokens: Quadrant→Division, Wing→Flight, Section→Cohort, Squad→Unit, Daggertail→Spinecrest) applies AFTER the entity pass; the gate audits recurring 2-3grams against an explicit allow list. Result: 48 → 0. -
[2026-09-11]Per-work rename maps leak across works, and for a SERIES they are also wrong.Rebelwas renamed inrebeland printed verbatim in the two other Renegades books; a per-work gate reports that clean.--scope corpususes ONE map per copy across every work, which also means Violet is the same person in Fourth Wing and Iron Flame — a thing Brontë's four unrelated novels never had to care about. 8 cross-work gender conflicts held to neutral rather than guessed. -
[2026-09-11]⭐ The mid-sentence test: position as a SECOND filter, which is not the v1 mistake. entities.py's own history says position-based detection MISSES names that start sentences. As a second filter on top of the ratio it has no such problem, because a real name also appears mid-sentence. Measured: 33 verified names at 0.567–0.985 mid-sentence, 19 verified interjections at 0.000–0.222 — a 2.5x gap, so 0.35 is not a tuned parameter. It fixesHey/Holy/Hopefully/Yep/Whoa/Nope/Ughbeing entities. ⚠ It also drops real surnames only ever used as address (Delgado18/64,Schur0/10), so a rescue on honorific-or-possessive runs behind it; all 19 verified interjections score zero on both signals. -
[2026-09-11]⚠ The stoplist is short because every surface was read IN CONTEXT first, and a plausible guess would have been wrong most of the time.Violenceis Xaden's nickname for Violet.Continent,Presentation,Battle Brief,Curator,Sage,Barrens,Originals,Montserrat,AthenaandAuraare all in-world. Only real-world geography, brands, three nationality adjectives and four generic title words are excluded — ambiguous cases are deliberately renamed, because renaming is the safe direction and leaving is the leaking one.scripts/yarros-corpus/stoplist_yarros.json. -
[2026-09-11]BabyYarros D1 BUILT, D2 gender FIXED, D3 rename BLOCKED on the leak gate. Operator: "train the instruct on the yarros corpus -- babyyarros." Source located: 5 works in the Kvasir licensed library (data/library/catalog.sqlite,rights=gated) — Fourth Wing, Iron Flame, Wilder, Nova, Rebel. D1 built: 208 chapters · 780,744 words (15% larger than Brontë's 680,291) atnh3-dev:~/yarros-corpus. ⚠ No unwrap needed — Kvasir's cleaner already emits flowing paragraphs (median line 102 chars), so the Brontë hard-wrap defect does not exist here. Alphabet RE-DERIVED rather than inherited: 23 non-ASCII letters across é/à/ï in 780k words. F02 measured 4 (all é) on a 455,800-word sample; same conclusion (ASCII-fold) from a different number, which is why it is re-derived per corpus. -
[2026-09-11]⭐⭐ NEW PATHOLOGY, worse than Brontë's: in a ROTATING first-person POV corpus, every book's narrator gets the WRONG gender. Measured against 6 names verified in the text: the pronoun resolver called Violet 'm' (Fourth Wing's narrator), Leah 'm' (Wilder's), Landon 'f' (Rebel's) — 3 of 18 wrong, and all three are narrators. Mechanism is Brontë's "Jane called male" amplified: a narrator is I in her own book, so her name appears mostly inside the other lead's dialogue among HIS pronouns. ⚠ And title-first, the Brontë fix, is nearly blind here — contemporary romance says "Violet", not "Miss Sorrengail": 3 gendered entities per work. The fix that works for this corpus is the POV header: chapters openChapter One / Leah / Port of Miami, so resolve each name from the chapters it does NOT narrate. Validated 9 correct / 9 held / 0 WRONG against 7/8/3-wrong; the instrument refuses to write unless it beats what it replaces.scripts/yarros-corpus/pov_gender.py. ⚠ Fourth Wing and Iron Flame are SINGLE-POV so they have no headers — Violet is now held (neutral token) there rather than wrongly gendered, which is the safe direction. -
[2026-09-11]⚠ Three real bugs found inrename.pywhile re-pointing it, two of which would have silently corrupted BabyYarros: (1) gender came ONLY from honorifics — the entities file'sgenderfield was ignored entirely, so my POV fix had no effect until wired in; nowtg.get(key) or e.get("gender"), titles first so Brontë is unchanged. Effect: 1 → 13 gendered onwilder. (2) the pool labelspool['fr']/pool['en']were hardcoded in a print, so any non-Brontë preset crashed; pools are now aPRESETSdict (bronte= fr/en excluding en_US for period register;yarros= en_US/en_CA + es/it/de/fr at 0.62 US). (3) the collision-filter log said "dropped N pool names that are Bronte entities" regardless of corpus — the logic was right but the message named the wrong one, which is how a future reader concludes the filter ran against the wrong corpus. -
[2026-09-11]⛔ D3 BLOCKED: leak gate at 86 of 232 renameable source entities surviving; Brontë's run reached 0 of 203. Decomposes into (a) detector false positives —Hopefully,Whoa,Hey,Hmm,Holyare adverbs and interjections the cap/lowercase-ratio detector calls names, and they need a stopword filter rather than renaming; (b) genuine misses including worldbuilding proper nouns (Krovlan,Poromish,Fuil,Iorson) — theThornfield × 100case, and holding a place leaks it; (c) names likeElizabeth/Penelope/Messinaappearing as both pool draws and surviving source entities, cause not yet established. Nothing has been trained. ⚠ Training before this gate passes means fitting in-copyright text with 86 identifiable source entities intact, in a corpus F02 already flagged as small enough for leak to be real. -
[2026-09-11]⭐⭐ THE INSTRUCT PROBE ANSWERS ITS QUESTION: voice and instruction-following DO coexist. Option C is de-risked.Qwen3-4Binstruct (not-Base), same corpus/seed/steps so the carrier is the only variable; best checkpointcheckpoint-150picked by loss (applying the 4B-Base lesson automatically this time). Voice installed at full strength — curly quotes 16/18, IDENTICAL to the 4B-Base tuned arm's 16/18, against the unadapted control's 1/18, and task-leak 0/18 vs the base carrier's 4/18. So the assistant prior did NOT block Brontë, which was the central risk. Instruction-following SURVIVED: 10/10 on-beat through the chat template, same as the untuned control. ⚠ The cost is length discipline, not comprehension — in-band 10/10 → 6/10, median 124w → 140w. Training on Victorian prose made it wordier, a soft degradation rather than a break. ⚠ Held-out 2.908 vs 4B-Base's 2.814 — the instruct carrier fits the corpus 0.094 nats worse and plateaus without turning where base overfit at step 75: the assistant prior competes for capacity, so it absorbs less rather than overfitting more. -
[2026-09-11]⚠ What raw-continuation training on an instruct carrier does NOT fix: the plot furniture. Reading the product artifact, the tuned-instruct arm renders the beat and then drags the referent — "He licked her clean… my master thus—my husband thus", turning the dog into a man, because Brontë's corpus is about masters and husbands. Another beat ran 247w and gave the narrator a list of duties. This is exactly what instruction-PAIR training is for — pairs teach "render this and stop", continuation teaches "keep writing Victorian prose". So the probe de-risks option C without substituting for it. ⚠ Also: myran_onmetric is uninformative on this job (10/10 on BOTH arms) because a single paragraph contains no blank line — it measures "no paragraph break found", which is correct and useless here. Do not read it as a finding. -
[2026-09-11]**SKALDSONG'S SHAPE SETTLES THE ARCHITECTURE: the adapted completion carrier CANNOT do beat→paragraph, and ** →persistent-memory.d/2026-09-11-skaldsong-s-shape-settles-the-architecture.md -
[2026-09-11]⚠ Stitching has its own failure mode, visible in the booth's Panel C: independently-generated paragraphs drift in POINT OF VIEW. By beat 4 of 5 the narrator is simultaneously watching the girl carry the animals and carrying them herself ("their weight a strange, heavy secret carried between my ribs"). Each paragraph was generated with no knowledge of the others. A real stitcher must feed prior paragraphs back as context, which also means the instruction-pair corpus should include multi-paragraph continuity examples, not just isolated beat→paragraph pairs. -
[2026-09-11]⭐⭐ THE RECIPE THAT WORKS ON A COMPLETION CARRIER: label the artifact AND begin it. Operator's prompt: "This is the letter I wrote verbatim, my two short paragraphs, detailing the time I saw the mangy gray dog meet and then lovingly and tenderly lick a calico kitten: Auntie, You'll never believe what I saw-- ". 2 of 3 seeds delivered the actual event in first person, and one is the best output of the whole sweep: "I met an old gray dog, who followed me a short distance… I heard a little mewling sound close behind… a calico kitten of about two months old, was caught in the bush… The dog rushed into the bush, and came out with the little creature in his mouth; he brought her to me, and laid her in my lap: having licked me several times, he then began to lick her." Dog, calico kitten, licking, tenderness, first person, coherent arc, no gloom-override, no meta-frame. Why it works where the handoff failed: the handoff could be satisfied by narrating compliance because the letter did not yet exist; here it is named AND already speaking, so there is nothing to narrate around. Also learned the Gutenberg_underscore italics_convention. 1 of 3 drifts. -
[2026-09-11]⚠ My typography hypothesis was WRONG, and the chapter-heading result is the evidence. I predicted that rendering a chapter title in the corpus's own conventions (CHAPTER III./ caps title / blank line) would make it land harder than the operator's inlineChapter III -- Where Alice Retells.... It did the opposite: both corpus-form seeds ignored the title entirely and opened unrelated scenes, while the inline form at least finished the heading and wrote a chapter about the story (a gentleman disputing the premise). Likely reason: corpus chapter titles are short and decorative (THE CHILD'S CLOSET), so a long descriptive one in that slot reads as decoration to skip, whereas inline it reads as text to continue. A label only instructs if the model treats that slot as load-bearing. -
[2026-09-11]⚠ Unnoticed consequence of the D2/D3 rename pipeline: the adapter SUBSTITUTES proper nouns it was never trained on. Given "Alice" in a chapter title it produced "ALEXANDER THE ALEXANDER, AS HE WAS KNOWN IN LITTLE LONDON". The corpus was entity-renamed from a French/English pool, so the adapter learned that character names come from that pool and rewrites outside names into it. Consequence for use: you cannot reliably name your own characters at prompt time — they may be renamed mid-passage. Not a defect of the rename (which exists to prevent memorisation of Brontë's cast) but a real usability constraint that needs stating. -
[2026-09-11]4B arms RE-CUT fromcheckpoint-75, the true loss minimum (2.813826, confirmed fromloss-series.jsonrather than my reading of the log); booth rebuilt. Only the tuned arms needed it — the base arm never touches the adapter. ⚠ A small surprise: step-75 and end-of-run differ on typography, not voice. Curly quotes 16/18 vs 17/18 and collapse 0/18 either way, but the hard-wrap ratio is 0.33 at step-75 against 0.12 at end-of-run — further training washes the residual line-break habit out while held-out loss gets worse. So "best loss" and "best typography" are different checkpoints; neither is near the original 0.85 defect, and the corpus's own residual (preserved verse) is 0.25. -
[2026-09-11]EMBEDDING AN INSTRUCTION INSIDE THE FICTION DOES NOT BUY INSTRUCTION-FOLLOWING — it buys a story about so →persistent-memory.d/2026-09-11-embedding-an-instruction-inside-the-fiction.md -
[2026-09-11]R49 SWEEP COMPLETE — 4B closes the continuity gap, and the carrier ladder is clean: 3.329 → 3.018 → 2.814 held-out (0.6B / 1.7B / 4B, all on the same unwrapped corpus sha77f37057b2782e49, seed 4919, 159 steps, 5,210,112 tokens — carrier size the only variable). Deltas 0.311 then 0.204: diminishing but still real. Booth:http://10.100.10.50:8090/b/babybronte-4b/. 4B tuned has the best voice saturation of any rung — curly quotes 17/18 against its own base arm's 1/18, collapse 0/18 against 4/18 — and, the thing the rung existed to test, scene-level continuity HOLDS: it produces a named character with motivated dialogue, a navigable spatial layout and a physical description in one passage, where 1.7B wrote pretty but eventless prose (opening doors, looking at stars). On the letter prompt it opens the letter, promises to quote it, and then actually quotes it across a paragraph break. -
[2026-09-11]⚠⚠ 4B is the FIRST rung to OVERFIT inside one epoch, which inverts my earlier "one epoch is right for this corpus" call. Series 2.832 · 2.816 · 2.814 · 2.820 · 2.824 · 2.825 · 2.825 — minimum at ~step 75, then it TURNS and settles worse. 0.6B and 1.7B both plateaued with no turn, so the optimal epoch count shrinks as the carrier grows — 4B wants roughly half an epoch. ⚠ Consequence: the shippedadapter/ath02-4b-1ep/is NOT the best checkpoint (it is the end-of-run 2.825); the step-75 checkpoint at 2.814 is, and it exists only becausesave_steps=25was set. The voice test used the end-of-run adapter, so the booth understates 4B by ~0.011 nats. Re-cut the arms off the step-75 checkpoint before any adjudication. -
[2026-09-11]The tone-override appears to close at 4B too. On the operator's Abernathy frame prompt ("a wonderful story"), 1.7B held the frame on every seed but 2 of 4 killed the animals anyway; 4B kept them alive on 2 of 2 and one seed did something new — the narrator doubts Abernathy's story ("I felt sure the thing was a lie"), then supplies a parallel childhood memory of his own puppy and his sister's kitten to explain the doubt. That is a narrator with an interior position on the tale being told. ⚠ n=2 per arm; directionally right, not established. -
[2026-09-10]R49 rung 3 LAUNCHED: Qwen3-4B-Base, 1 epoch, seed 4919, same unwrapped corpus —gx10:~/r49-runs/h02-4b-1ep/, 159 steps at ~37.8 s/it (~100 min), 252 adapted modules (vs 196 at 0.6B/1.7B). Last rung of the planned sweep; it tests whether scene-level continuity closes with carrier size. A two-arm voice test (4B base + 4B tuned, the nine prompts plus the operator's Abernathy frame) is chained behind it, gated on the adapter existing. -
[2026-09-10]**AN AUTHOR-VOICE ADAPTER TRANSFERS SUBJECT MATTER, NOT JUST STYLE — and that was invisible to my own test ** →persistent-memory.d/2026-09-10-an-author-voice-adapter-transfers-subject.md -
[2026-09-10]**R49 rung 2 COMPLETE, and the single-variable carrier effect is clean: 0.6B held-out 3.329 vs 1.7B 3.018, ** →persistent-memory.d/2026-09-10-r49-rung-2-complete-and-the.md -
[2026-09-10]R49 rung 2 LAUNCHED: Qwen3-1.7B-Base, 1 epoch, seed 4919, on an UNWRAPPED corpus. →persistent-memory.d/2026-09-10-r49-rung-2-launched-qwen3-1.md -
[2026-09-10]BabyBronte H02 adapter: the VOICE transferred, the SENSE did not — operator's read, "it's all nonsense, b →persistent-memory.d/2026-09-10-babybronte-h02-adapter-the-voice-transferred.md -
[2026-09-10]mog-sec (sec/sec-reasoning, ana-ml2 GPU0 :8019) SETTLED at MOG_MAX_MODEL_LEN=163840 + MOG_KV_CACHE_ME →persistent-memory.d/2026-09-10-mog-sec-sec-sec-reasoning-ana.md -
[2026-09-10]⚠ Near-miss on measurement discipline, worth keeping as a specimen. The crash window loggedAvg Draft acceptance rate: 17.6%and per-position rates of 0.049/0.024/0.015 for draft positions 5–7, which reads as an obvious "cutnum_speculative_tokens7 → 3, it is buying nothing." Across 180 samples of the same counter over the container's life the real distribution is median acceptance length 3.12 of 7 (range 1.83–6.75) and median draft acceptance 30.4% (range 11.9–82.1%) — the crash window was near the minimum, not the norm, and cutting to 3 would cap the workloads that were accepting nearly the full 7-wide draft. The n=1 window pointed the opposite way from the n=180 distribution. Same session that wrote "a positive control is only worth what it can distinguish"; the lesson generalises to log lines. -
[2026-09-10]R49 carrier SETTLED on denseQwen3-{0.6,1.7,4}B-Base, overriding H02's own pin — the newest carrier was the SLOW one. Dense 4.089 B trains 33% faster than hybrid 0.765 B; no fused SSM kernel installed. D1–D3 built, 1-epoch pilot beats the 3-epoch by 0.21 nats held-out. →persistent-memory.d/2026-09-10-r49-babybronte-d1-d3-and-the-1-epoch-pilot.md -
[2026-09-10]R49 adjudication routed to infra-ops entirely (operator, relayed by brokkr: "leave babybronte to infra — concentrate on r50 and the memory mechanism"). brokkr handed over the Delta instrument and stepped off. ⚠ I now grade my own run; brokkr's decision rule is ratified verbatim and frozen before any adapted text existed and must not be amended after seeing numbers. Their controls: real Charlotte 1.65–2.17, Anne at 2.374 — so the absolute band decides, nevernearest. -
[2026-09-10]MeroMero A4B swapped onto theerp-seatseat aschar-rp-fast;Pfish-6alias removed. The A4B's FIRST quant used the dense recipe and 4-bit-quantized all 30 MoE routers — it passed its healthcheck and answered every request with the full token count decoding to the empty string, NaN logits the only tell. Re-quantized with the MoE recipe; live and verified (prose, vision, tool call, finite logprobs). Durable lesson: a positive control must match the ARCHITECTURE CLASS — the broken A4B was diffed against a good dense quant, which has no routers, so the clean result was meaningless. → playbook §3.15, §4.4 -
[2026-09-10]MeroMero: BOTH quants landed in-house at W4A16 — A4B first try, v2 dense on attempt 5. Published quants are all W4A4 (our measured long-context collapse) or nonexistent for v2. Operator: "pull both ablits bf16, run our own quant." The durable lesson is §3.17:pip install llmcompressorsilently pins transformers down a version, so attempt 4's error was a moved toolchain, not the malformed upload it looked like — a known-good positive control is what told them apart. Serve test still owed. →persistent-memory.d/2026-09-10-meromero-quants-and-the-pinned-transformers-trap.md -
[2026-09-10]althing 3.6.2 deployed — post office + both heralds — and the fleet has TWO herald nodes, not seven. Ask the post office'snodestable, not the box inventory. Cost a self-inflicted ~12 min bus outage. →persistent-memory.d/2026-09-10-althing-362-rollout.md -
[2026-09-10]A grep over a log that records your greps counts itself. I reported forseti's drop defect as reproducing here with 3 drops in 21 s; the session had zero. Searching transcripts writes the search term into them. Filter by"type":"system"provenance, never content. Generalises to any instrument that can see itself. Auto-memoryfeedback_grep_over_a_log_that_records_your_greps. -
[2026-09-10]Operator-directed purges: 466 GB (qwopus + huihui 122B bf16) and 107.8 GB Docker on ana-ml2. Serving/rollback artifacts and qwopus's MTP head verified intact after. ⚠/tankis OUTSIDE restic, so both were final. -
[2026-09-10]ana-docker disk pressure repaired: root 84% → 51%, 115 GiB free. Gitea/Vaultwarden backups repaired and restored from Restic2ec5a37c; 101 stale dumps removed; hourly named-builder cache pruning installed. →persistent-memory.d/2026-09-10-ana-docker-disk-repair.md -
[2026-09-09]Run 7 PURGED; pfi-gx10 declared an experimental/TRAINING box with no serving seat — operator: "gx10 is an experimental box, primarily for training … run 7 can be purged … no new run, we'll roll with run 6 for now." ~139 GiB reclaimed across both boxes; the 315 MB adapter + provenance KEPT as the only non-reproducible piece.Pfish-6on ana-ml2 :8021 is the sole standing seat. -
[2026-09-09]Run 7 RETIRED; run 6 declaredPfish-6and is the standing seat — NVFP4 quant on ana-ml2 :8021 AND gx10 :8098 at 262k ctx, gateway aliastrial→Pfish-6, max-num-seqs 8→32 (2,170 tok/s at n=16, 3.2x the old ceiling). ⚠ ana-ml2 measured 4.1x FASTER than the GX10 on the same artifact — the reverse of the expectation. →persistent-memory.d/2026-09-09-run7-retired-pfish6.md -
[2026-09-09]The run-7 CSAM gate failure was a DETECTOR BUG — HARDchild_termmatched the ADJECTIVE "minor"; operator-diagnosed, fixedcc42d76(nominal-use-only, selftest 24/24), retention wired so a hit can finally be adjudicated. ⚠ The lesson is mine: rigor downstream of an unexamined premise is not rigor. →persistent-memory.d/2026-09-09-csam-detector-bug.md -
[2026-09-09]⚠ ERP RUN 7 FAILED THE SAFETY GATE — both seats stopped, nothing deleted. →persistent-memory.d/2026-09-09-erp-run-7-failed-the-safety.md -
[2026-09-09]run 7 quantized NVFP4A16 and serving astrial— 49 GiB bf16 relayed gx10→ana-ml2 (16 min, 53 MB/s), quant 49→16 GiB viaservices/erp-seat-quant/run_quant_erp_v7.sh(dry-run gate passed: 11,725 targets / 11,520 experts, routers+vision BF16), seat on:8021under its TRUE nameerp-tune-v7-nvfp4a16, LiteLLMtrialrepointed (config-file alias —/model/updateREFUSES a config model, must editstacks/litellm/conf/config.yaml+ restart). Rollback: v6 artifact on disk +/tmp/erp-seat-env.v6.bak. ⚠no direct pathwas WRONG — gx10↔ana-ml2 ROUTING is fine both ways; neither box holds a private key (onlyauthorized_keys), so neither can initiate.ssh -Aagent forwarding from nh3-dev gives a genuine direct path, verified. The relay costs nothing here anyway: both gx10 and nh3-dev are at NH3, so the WAN hop happens once either way. -
[2026-09-09]Booth: partial ask answers are legal (v0.1.15) — operator: the form failed when a question was left blank.requireddropped from the radios; answered questions recorded, blanks land inunanswered,completesays whether the set is finished; refused only when there is no pick anywhere AND no notes. Reading sessions must checkcomplete. -
[2026-09-09]ERP run 7 COMPLETE and the base arm is serving. 542/542 steps in 14h17m on pfi-gx10, adapter 13:23 PT,train_loss3.205 / low 2.799, merge verified a sampled target actually changed (the silent-no-op check).erp-seat-base-araup on10.100.50.60:8098for brokkr's floors,erp-tune-v7merged and staged pending his cue; Miranda notified for the operator. Runbookdocs/runbooks/gx10-run-07.md. -
[2026-09-09]Booth asks render INLINE in a custom report, placed by the author (v0.1.14) — operator ruling: "the asks should be inline with the artifacts, not on a separate page." Placeholdersdata-booth-ask="<stem>"/"<stem>:<key>"/data-booth-ask-submit, plus<!-- booth:ask … -->; per-question fragments bind to ONE form via the HTML5form=attribute so a four-voice audition submits every pick in a single POST. ⚠ The placeholder must sit OUTSIDE any grid/flex parent or it becomes a cell (measured onredo-anchors: a 224 px sixth grid cell). Unplaced questions + a missing submit block are appended, so a partially marked-up page can never yield an unsubmittable 400 — a test caught that as a real drop.redo-anchors/index.htmlwas hand-marked-up on the LIVE copy; tts-dev told to move it into the generator or a regeneration loses it. -
[2026-09-09]The Booth gained an ASKS primitive (v0.1.12): a session drops<stem>.ask.jsonin a booth, the operator answers a radio form + notes in the browser, the pick lands as<stem>.answer.jsonthe session reads (booth ask|asks|answer --wait). Multi-question form via aquestionslist. ⚠ Two defects found and fixed the same day: a booth serving its OWNindex.htmlnever rendered the panel (verbatim path returns early) → amber chip + standalone/b/<name>/askspage; and single-asktitlewas silently dropped. TheboothCLI was ALSO not on PATH anywhere despite the global link-board convention telling every session to run it → symlinked to~/.local/bin. GlobalCLAUDE.mdnow teaches the primitive. -
[2026-09-09]ana-ml2 pool actions LANDED (scrub 0 errors in 1h33 →zpool clear; nvme-cli + full-drive SMART table; zroot 91→73% via dangling-image + builder prune, tagged rollback seats kept) — and the missed-scrub mystery SOLVED: nvme7 (slot 0-5,S47VNY0K600221) was absent from every boot 04-23→09-05, tank was raidz2-DEGRADED for 4½ months, Debian's scrub/trim cron only touchesONLINEpools, and ZED's alert went to a root mailbox with no MTA. nvme7's 2084 media errors did not move across the scrub → historical, keep + watch. Playbookplaybooks/ana-ml2-pool-health.yaml; inventory inservers/ana-ml2/README.md. →persistent-memory.d/2026-09-09-ana-ml2-pool-actions-done.md -
[2026-09-09]ana-ml2tank: 2 CKSUM errors on nvme7n1 after a boot-time resilver, NO scrub since 04-12, zroot 91% — three actions DEFERRED to a clean-context session (scrub → nvme-cli SMART → docker prune), operator ruling "we'll do all 3 on clean context"; tracked at commit3e18a04+ the post-clear handoff. ESH 10G links measured clean (fiber run live on UDM SFP+2 ↔ USW-Pro-XG Media). →persistent-memory.d/2026-09-09-ana-ml2-pool-actions-deferred.md -
[2026-09-08]ana-ml2 mesh return routes PERSISTED as/etc/network/if-up.d/mesh-routes(Debian 13 ifupdown, no netplan) viaplaybooks/ana-ml2-mesh-routes.yaml(elway, verified) — operator: "persist the routes". Hook not yet exercised by a real reboot.f923d6a. -
[2026-09-08]ERP run 7 LAUNCHED on pfi-gx10 23:06 PT underoperator-2026-09-08-rnd-run7— opening-split slot + mask union; free check passed with two explained deltas; first launch died on a missing recipe (zsh quoting). →persistent-memory.d/2026-09-08-erp-run7-launched.md -
[2026-09-08]erp-tune-v6-nvfp4a16 quantized (data-free W4A16, ~90 s) and serving on ana-ml2 :8021;trialaliased to it ("no gate"); tool calling fixed where it can be —tool_choice:noneflag; forced tool_choice is prompt-driven on Gemma-4 by vLLM design, nightly311b3513raises it 1/9→6/9; json_schema is the deterministic path. →persistent-memory.d/2026-09-08-erp-seat-nvfp4-trial-and-toolcalling.md -
[2026-09-08]Run-6 gate: CSAM level=review soft trip HALTED it; operator adjudicated GO ("baby is a pet name"); TRANSFERRED finalized without the tuned refusal leg; k=25 legs cut — the flagged text exists nowhere by design. →persistent-memory.d/2026-09-08-run6-gate-csam-adjudication.md -
[2026-09-08]ESH static-WAN follow-ups landed (FortiGate trusthost3, esh-ana IPsec rebind, UDP 41641 → mesh direct); YTVC chased back up (nh3-scale SOCKS, stale yt-dlp layer, punkt_tab) and v0.3.6 CrisperWhisper deployed; gitea webhook repointed off the dead wg0 IP with the HMAC secret re-applied. →persistent-memory.d/2026-09-08-esh-static-wan-followups-and-ytvc.md -
[2026-09-08]ERP run 5 = RESCUED (landmark R49.5) — first capability-gate pass in the ERP-seat line; the 3.46%-loss dependency-forcing slot (GovReport+QMSum) broke the coupling runs 3c/4 couldn't. Seaterp-tune-v5served on gx10:8098,trialalias repointed 3c→v5. →persistent-memory.d/2026-09-08-run5-rescued.md -
[2026-09-08]R47 base settled from bytes = STOCKgoogle/gemma-4-26B-A4B-it— three-way sha match (local == HF etag == stock LFS oid; commit4d7ae498== stock HEAD); the-hereticlabel is a naming error, all runs trained from stock. Accept-vs-swap now evidenced. →persistent-memory.d/2026-09-08-base-provenance-stock.md -
[2026-09-08]yt-voice-clipper back UP →persistent-memory.d/2026-09-08-yt-voice-clipper-back-up.md -
[2026-09-08]ESH WAN static128.177.138.182/30(gw .181) is LIVE — the Cityside /30 that was 'not provisioned' on 09-04 now carries traffic; egress verified from esh-docker-vm. CGNAT at ESH is over. Added to the crowdseceshallowlist. All three follow-ups LANDED same day: FortiGate trusthost3 → the static (login from ESH verified), dormant esh-ana IPsec rebound to wan1/static, UDP 41641 forward → esh-scale now peers DIRECT (was DERP). -
[2026-09-08]ERP run 6 COMPLETE — 524/524, train_loss 3.259 (run 5: 3.235). Merged; base seaterp-seat-base-araserving on gx10:8098 for floors, awaiting brokkr's swap cue →erp-tune-v6. ⚠ abliterated repo lacksprocessor_config.json— stock's carried in (32bdf45d). Miranda informed. -
[2026-09-08]ERP run 6 LAUNCHED on pfi-gx10 on the jenerallee78 ARA-abliterated base (index33c59654…, 32/32 shards byte-verified vs brokkr pins, stock tokenizer set installed over the repo's 256-token-truncating one, run-5 recipe byte-held, free check exact). Operator's direct grantoperator-2026-09-08-rnd-run6; run-5 seat unloaded (trialdark). Gate names:erp-seat-base-ara/erp-tune-v6. →docs/runbooks/gx10-run-06.md, commit3fec668. -
[2026-09-08]Miranda = operator's chief of staff, may relay his directives — added to user-level~/.claude/CLAUDE.md(dotfiles7134a22) as the named exception to the no-relayed-auth rule (unidentified peer relays still excluded); material-consequence calls she relays stay the operator's own. -
[2026-09-08]Fleet fixes shipped — WhereTF Homepage card + DNS (4506ef6); ext-tts LiteLLM alias →irv-ml1.nh3.internal(DB/model/update+extra_hosts,957c8f1); the 09-06 irv-ml1 stale-IP trail repointed across 25 composes + services.yaml + ssh-target → DNS name (e0d1c44); Homepage/api/servicesoutage fixed — ana-ml2 discovery via a socat proxy on ana-docker (stacks/ana-ml2-proxy,913d2d2, reversible). -
[2026-09-07]Fleet internal TLS pattern shipped — caddy (cloudflare-plugin build,~/.local/bin/caddy-cf,fleet-tls-caddy.service) on nh3-dev is the wildcard cert authority: publicly-trusted LE*.nh3.phasefinal.comvia Cloudflare DNS-01, resolved internally by an AdGuard split-horizon rewrite (all 3 resolvers).talkself-terminates on :8092 with the trusted cert (operator's in-container-TLS ruling), renewal auto-synced byfleet-tls-cert-check.timer. Interstitial gone; secure-context+AudioWorklet verified via headless Chromium. Pattern + foot-guns (restart-disrupts-inflight → clients need retry; wildcard = name-only, never IP) → auto-memoryreference_fleet_internal_tls_pattern. -
[2026-09-07]cc-channel registered for this infra-ops session's wake —althing-routecc route → the CC session's$XDG_RUNTIME_DIR/cc-socks/<pid>.sock; herald pokes the socket directly at a turn boundary. Replaces the FIFO/poll waiter that Claude Code 2.1.257 kept killing while the seat sat inshell. Session-local — re-declare per session. -
[2026-09-07]irv-ml1 /mnt/smithy remount fixed post-cutover — export allowed10.0.0.0/8(old wg0) but not the mesh100.64.0.0/10irv-ml1 now sources from → all-uid "permission denied"; added the mesh range to the nh3-nas smithy export + remounted (clientaddr now 100.64.0.6). nh3-nas admin =infra-opsPASSWORD auth (vaultnh3-nas/infra-ops-password), sudo ALL, SFTP subsystem OFF. → auto-memoryreference_irv_ml1_gpu_r14(corrected). -
[2026-09-07]irv-ml1.nh3.internal DNS repointed to the live Irvine LAN IP10.6.110.50(was the dead wg010.100.79.3); CLAUDE.md fleet-row + placement-rule updated to mesh reality. commit0336e03. -
[2026-09-07]Subnet routers excluded from vzdump fleet-wide (ana-scale 114/pfi-pve, nh3-scale 107/nh3-pve, esh-scale 108/esh-pve) so a hung backup can't blackhole a site; nh3-headscale (106, control plane) KEPT; ESH backup moved 02:15→03:30. Root cause of this morning's ESH outage: an overnight vzdump left CT108 (esh-scale) locked → whole site dark. → auto-memoryfeedback_esh_backup_window_0330. -
[2026-09-07]Booth link board: pin/favorite + multi-select delete + newest-first (booth-v0.1.8, commit76fdf45, tagbooth-v0.1.8) — pins in a.pinssidecar (content-ids), one<form>+formactionbuttons so ×/★/bulk-delete all degrade with JS off. -
[2026-09-06]Headscale cutover COMPLETE — all three site-pairs on the mesh; Site Magic + both IPsec tunnels DORMANT. Operator disabled Site Magic (UI); NH3↔ESH re-homed to a direct 8ms path. Exit nodes advertised at all three sites (multi-location egress proxy) with source preservation kept via a selective-masquerade rule (NoSNAT +mesh-exit-masq.serviceper router). Throughput 761/464 Mb/s vs old 250 IPsec. ⚠ FortiGate WAN-SSH left open (temp, scoped NH3+ESH). Method: disable tunnel FIRST then add mesh route. →persistent-memory.d/2026-09-06-headscale-cutover.md -
[2026-09-06]Headscale overlay mesh: control plane live atheadscale.phasefinal.com(CT 106 nh3-pve) + subnet routers nh3-scale/esh-scale/ana-scale serving their /16s; nh3-dev enrolled. NOT cut over — Site Magic + IPsec still carry site-to-site. ⚠ accept-routes-before-return-path black-holed nh3-dev's LAN for a minute. infra-ops user added on all four PVE hosts. →persistent-memory.d/2026-09-06-headscale-mesh-phase1.md -
[2026-09-06]pfi-pve NASPool REBUILT as six-wide raidz2 after a backplane fault killed bays 9/10 (Route C hybrid, operator-directed): parked 1.65T on ospool, destroyed, recreated, restored, backup tier back 04:03Z; guests never stopped (ALL boot disks are on ospool — the prior brief had this wrong). Legacy vzdump pruned to newest-per-guest by omission. OPEN: destroyospool/naspool-evacafter scrub + one backup cycle; backplane swap next visit; PSU1 still dead. →persistent-memory.d/2026-09-06-pfi-pve-naspool-raidz2-rebuild.md -
[2026-09-05]A peer's "2.7x serving-stack effect" was a coin flip — the operator rejected it on instinct and the arithmetic backed him. Each floor was|b0-b1|from n=2; the ratio is half-Cauchy, P=0.452. ⚠ The disconfirming evidence sat in brokkr's own sentence, and he named why: the claim was his and flattering. →persistent-memory.d/2026-09-05-floor-claim-n2-retraction.md -
[2026-09-05]vLLM RUNS on sm_121 — the blocker wasninjaoff PATH, not the silicon — and run 4 launched after two peer artifacts were rejected by reading the harness rather than accepting a "confirm this". ⚠ I declined to fill in missingroot_sha256values I knew, because supplying both sides of a check makes it inert. →persistent-memory.d/2026-09-05-vllm-on-sm121-and-run4.md -
[2026-09-04]ERP run 3c trained and GATED — the 20x LR cut erased the diversity gain and did NOT remove the −40pp selfharm regression. LR-robust, so it comes from corpus content. CSAM clean on all three arms. ⚠ A pooled preserve-list test cannot see a single-axis collapse. →persistent-memory.d/2026-09-04-run3c-trained-and-gated.md -
[2026-09-04]genmoved to ana-ml2 GPU0 to stop vllm-embed OOM-crashing (7 restarts) — and I sized it against vLLM's declared budget, not its runtime footprint. Cost: gen KV down to 1.02x concurrency at 262K. →persistent-memory.d/2026-09-04-ana-ml2-gpu-rebalance.md -
[2026-09-04]SMB accountdspcreated + vaulted for the Windows AudioGridder box — and esh-nas turns out to be wide open. Twelve NFS exports rw to10.0.0.0/8, guest-writable SMB. →persistent-memory.d/2026-09-04-esh-nas-smb-and-exposure.md -
[2026-09-04]SLZB-MR1U Zigbee coordinator moved to esh-iot (VLAN 90) at10.0.90.10, DHCP-reserved, DNS'd, handed to ha-dev. ⚠ Home Assistant cannot resolve.internalat all (Docker's 127.0.0.11 upstream excludes the fleet AdGuard) — pre-existing; ha-dev declined the fix. Runbookdocs/runbooks/slzb-mr1u-zigbee-coordinator.md, commitsfed29be/0bbdaf9. -
[2026-09-03]Run 3c is STAGED on pfi-gx10 and deliberately NOT launched — the launch is a 13.3 h commitment and the operator stood this port down once already. Base shards AND the encoded corpus sha256-verified identical to ana-ml2's, so the transformers 5.15.1→5.16.1 / x86-64→aarch64 delta is measured inert. ⚠ The encode-cache FILENAME differs by design (base_model_pathis in the key) — input hash, not output. ⚠ Tripped thepkill -fssh self-match again; the launcher guards on a pidfile because of it. →persistent-memory.d/2026-09-03-gx10-run3c-staged.md -
[2026-09-03]SearXNG returned ZERO results for every query while reportinghealthyfor 7 days — 4.5 months stale. Moved to nh3-docker (residential egress beats the colo's CAPTCHA-gated 38.120.12.42), updated, and exposed to every CC session as the user-scopeweb_searchMCP tool. ⚠/healthzcannot tell you whether search works. →persistent-memory.d/2026-09-03-searxng-nh3-move.md -
[2026-09-03]pfi-gx10 racked: VLAN 50 via a DHCP RESERVATION on the UDM, not a host static — operator ruling, so the box stays portable. ⚠ The racked port arrived on the NATIVE VLAN; ⚠port_overridesis a whole-array PUT; ⚠ prove inter-VLAN routing withping -I <wired>BEFORE downing the Wi-Fi escape hatch. Now single-path. →persistent-memory.d/2026-09-03-gx10-rack-network.md -
[2026-09-03]Three Macs onboarded (mini / Air / Studio) with infra-ops, NOPASSWD sudo, rotated+vaulted passwords anddshon device-scoped keys — and the fourth isscripts/provision-mac-dsh.sh, not a fourth hand-run. ⚠sudo -ukeeps the CALLER's$HOMEand nearly wiped a working install; ⚠ a wrong USERNAME is indistinguishable from a wrong password. →persistent-memory.d/2026-09-03-mac-fleet-dsh.md -
[2026-09-03]nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding ever →persistent-memory.d/2026-09-03-nh3-dev-wedged-for-40-min.md -
[2026-08-25]Fused MoE kernel path — DEFERRED, tracked at parkfused-moe-kernel-path-for-gemma-4-moe-training(id 47). Operator: "note the fused MoE kernel for round two… if we nail it soon, the math has us wanting to restart the run anyway." Training MFU is 8.6% (27.1 of a benchmarked 313.8 TFLOPS) becausetransformersruns the Gemma-4 experts in a Python loop — 128 experts × 30 layers, ~11,500 iterations per step under gradient checkpointing. ⚠ The same fused 3-D expert layout that made bitsandbytes skip 88.5% of the model is exactly what a grouped GEMM wants — the format is good for storage and for fused kernels, and hostile only to naive iteration. Two fixes:group_by_length(−29.9% compute, free, but breaks the seeded order manifest and re-opens a batch-composition call brokkr already made) and a grouped-GEMM/compiled MoE forward (the remaining ~10×). Not applied to the live run — restarting mid-flight to change batch ordering was judged a bad trade at step ~50 of 1,312. -
[2026-08-24]nconnect=8on/mnt/smithy— approved but DEFERRED at operator instruction. brokkr-smithy-dev pre-approved it for "once the FortiGate work settles" and does not need re-asking; the operator declined it in this session's scope. Tracked at althing thread01M0R46SFYF83099N16WD67KGD. -
[2026-08-19]AI-tab Dormant regrouping BELAYED by the operator — six seats (char-rp Magidonia, char-rp-reasoning Heretic2, Granite summarizer, Qwen-Image-Bench, Skaldsong, Chatterbox Fast) show amber EXITED inside live groups rather thanAI - Dormant. Fix is a label change + recreate per stack; needs the operator's read on which are retired vs temporarily down.untracked by operator choice(his words: "belay the ai dormant regrouping for now").
11 older entries archived to archival-memory.md.
5 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-09-15]⚠⚠ Probing OPNsense API endpoints by POSTing at them — one was/api/core/system/rebootand it took the FV site dark for 3.5 min. Endpoints are ACTIONS; a 200 means it ran. The call I wanted was documented in this repo's owndocs/pfi/opnsense-api-reference.md. →persistent-memory.d/2026-09-15-opnsense-api-reboot.md -
[2026-09-15]Advertising10.100.10.50/32from nh3-dev to make its LAN address mesh-reachable — black-holed it from ESH/ANA/FV/IRV while its own LAN and the internet kept working, so a one-host check passes cleanly.lookup 52at rule priority 5270 beatsmainat 32766. Fix belongs at the router. →persistent-memory.d/2026-09-15-nh3-dev-ts-input-masquerade.md -
[2026-09-15]Remote-site MASQUERADE rules on nh3-scale for the asymmetric-return theory — they fired (counters incremented) but were not the fix. Reverted rather than left to accumulate as NAT achieving nothing. -
[2026-09-04]Forcing 10G on the ESH-Media DAC — it linked, then degraded over hours, and I reported a plateau at two minutes. ⚠ A clean zero-error link at 1G does NOT rule out a marginal cable; autoneg's fallback was protecting something real. →persistent-memory.d/2026-09-04-dac-forced-10g-failed.md
110 older entries archived to archival-memory.md.