memory: snapshot — lv-mccarthy training launched on gx10, and the next voice seat is measured rather than chosen
In-flight rewritten to the live training run (~150/1380, ETA ~00:45 PT) with the --save-total-limit finding that would otherwise have deleted the epoch-1/epoch-2 checkpoints both prior gates were decided on. Two decisions added: the next-seat ranking (Faulkner, Morrison, Chandler -- and the finding that the corpus size ranking inverts the voice ranking, with King and Christie as the two biggest non-candidates), and the romantasy register measured on the gate's own char-bigram instrument (Yarros is the cluster outlier we already shipped; Maas is the centroid and so the worst pick; Kenyon at 27 val units if the lane gets a seat). Auto-archival: 4 entries moved to archival-memory.md; 4 held back by the open-deferred guard.
This commit is contained in:
@@ -4,6 +4,181 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re
|
||||
|
||||
## Recent decisions (archived)
|
||||
|
||||
- `[2026-09-03]` **SearXNG was returning ZERO results for every query while reporting `healthy` — moved to nh3-docker, updated, and exposed to every CC session as an MCP tool.**
|
||||
|
||||
**The failure.** The ana-docker instance answered `/healthz` every 30s, showed
|
||||
`Up 7 days (healthy)` with 0 restarts, and had a green Homepage card — while
|
||||
returning **0 results for every query tested**. It was running **2026.4.17
|
||||
against a current 2026.9.3**: 4.5 months of engine scrapers rotting against
|
||||
sites that had changed their markup. SearXNG ships near-daily releases for
|
||||
exactly this reason.
|
||||
|
||||
⚠ **`:latest` means "latest AT PULL TIME".** Nothing re-pulls on its own. A
|
||||
container created in April on `:latest` is pinned to April forever.
|
||||
|
||||
⚠ **`/healthz` proves the web app answers and says NOTHING about whether search
|
||||
works.** That is the whole lesson. Same shape as the nh3-dev "failing disk" that
|
||||
was a stalled backup, the runner audit that trusted liveness for identity, and
|
||||
the statusline bell that measured a mechanism.
|
||||
|
||||
**Proven before acting**: the new image, same settings file, same host, same
|
||||
query, in a throwaway container → **20 results where the running one returned
|
||||
0**. Network was ruled out first — from inside the container DNS resolved and
|
||||
mojeek/wikipedia were reachable, so engines were reachable and the parsers were
|
||||
the broken part.
|
||||
|
||||
**Why NH3 and not an in-place update** (operator's call, and the measurement
|
||||
backs it):
|
||||
|
||||
ana-docker egress 38.120.12.42 datacenter -> DuckDuckGo/Startpage CAPTCHA
|
||||
nh3-docker egress 70.230.226.88 residential -> not gated the same way
|
||||
|
||||
Search engines gate datacenter ranges. Same reason the fleet keeps a residential
|
||||
SOCKS5 proxy on nh3-dev for yt-dlp — applied at the source instead of proxied
|
||||
around. `outgoing.proxies` has the fallback commented in place if NH3's egress
|
||||
ever changes.
|
||||
|
||||
⚠ **Not a complete fix.** `brave`, `duckduckgo`, `startpage` still CAPTCHA from
|
||||
NH3. `google cse` carries general search at ~20 results/query; `yandex`, `wiby`,
|
||||
`github`, `stackoverflow`, `marginalia` work. **General search is effectively
|
||||
single-engine** — if google cse breaks, it goes quiet again.
|
||||
|
||||
**Two config defects, both silent:** `base_url` still named
|
||||
`searxng.pfi.local`, retired 2026-08-19, while the env said otherwise (env wins,
|
||||
so nothing broke and the file lied to every reader); and the
|
||||
`karmasearch.videos` removal key never matched because the engine's real name
|
||||
has a space in it.
|
||||
|
||||
**`scripts/searxng-health.sh` asserts results > 0** across three unrelated
|
||||
queries. That is the only check that could have caught this — the mechanism was
|
||||
healthy throughout.
|
||||
|
||||
**The MCP tool** — `services/searxng-mcp`, `uv tool install`, registered
|
||||
`claude mcp add --scope user searxng searxng-mcp`, so every CC session gets
|
||||
`web_search`. ⚠ Zero results **raise** rather than returning an empty list: an
|
||||
empty list is indistinguishable from a broken aggregator, which is precisely how
|
||||
this hid. Same principle as althing's "unreachable post office is an OUTAGE,
|
||||
never an empty inbox".
|
||||
|
||||
⚠ Written against **mcp 2.x** (`FastMCP` → `MCPServer`; the v1
|
||||
`@app.list_tools()` decorator is gone and fails at import). ⚠ **`uv tool install
|
||||
--force` served a CACHED build** and silently reinstalled the old code — the
|
||||
installed file still had the v1 API after the source no longer did.
|
||||
`--reinstall --no-cache` fixed it; `md5sum` of source vs installed is what
|
||||
caught it.
|
||||
|
||||
Old instance stopped and removed; DNS alias repointed to
|
||||
`searxng.nh3.internal` → 10.100.50.40. Secret vaulted at
|
||||
`nh3-docker/searxng-secret`. Commit `0f748ea`. See [[2026-09-03-gx10-rack-network]]
|
||||
for the other UniFi-side change the same day.
|
||||
_Archived 2026-09-18._
|
||||
|
||||
- `[2026-09-03]` **pfi-gx10 racked and networked: VLAN 50 via a DHCP RESERVATION, not a host static; Wi-Fi down.**
|
||||
|
||||
`pfi-gx10.nh3.internal` → **10.100.50.60**, wired only.
|
||||
|
||||
**Operator ruling, and the better design:** put the address on the
|
||||
**switch/firewall side** as a DHCP reservation and leave the host on DHCP. A
|
||||
host-side static works until the box moves, and then it is a stale netplan file
|
||||
on a machine whose address you no longer know. A reservation moves with the MAC.
|
||||
|
||||
UniFi switch port 22 native network -> nh3-servers (VLAN 50)
|
||||
UniFi client reservation -> 30:c5:99:3d:a7:45 = 10.100.50.60
|
||||
host unchanged, still DHCP
|
||||
|
||||
`playbooks/gx10-rack-network.yaml` was pre-written to apply a **host static** and
|
||||
was NOT used — annotated as retired at its top. Its safety *ordering* was
|
||||
followed and is still right.
|
||||
|
||||
⚠ **The port arrived on the native VLAN**, not the server VLAN — it DHCP'd
|
||||
`10.100.0.111` from `nh3-default`. The switch port had to be repointed before
|
||||
anything else could work. Do not assume a racked port is on the VLAN you asked
|
||||
for.
|
||||
|
||||
⚠ **`port_overrides` is a WHOLE-ARRAY PUT.** Anything omitted is deleted. Two
|
||||
unrelated overrides (ports 21, 23) were read, backed up to a file, preserved and
|
||||
written back.
|
||||
|
||||
⚠ **The step that is easy to skip and expensive to miss:** while Wi-Fi was still
|
||||
up, traffic from the box to nh3-dev **preferred `wlP9s9`** — that interface sits
|
||||
directly on the userland subnet — so "I can reach it on the new address" proved
|
||||
NOTHING about the wired path. Downing Wi-Fi on that evidence is a coin flip on
|
||||
inter-VLAN routing, and losing it is a rack visit. Forcing the interface is what
|
||||
settled it:
|
||||
|
||||
ping -c3 -I enP7s7 10.100.10.50 0% loss VLAN 50 -> VLAN 10
|
||||
ping -c2 -I enP7s7 1.1.1.1 0% loss egress
|
||||
|
||||
Only then did Wi-Fi come down, as its own step, `/etc/netplan` backed up to
|
||||
`/etc/netplan.bak-preWifiDown`. `nmcli radio wifi off` persists across reboot —
|
||||
verified by reading `/var/lib/NetworkManager/NetworkManager.state` back.
|
||||
|
||||
⚠ **The box now has exactly ONE path.** If that switch port or the reservation
|
||||
breaks it is a rack visit; the escape hatch is deliberately gone. Correct end
|
||||
state for a racked server, but a posture change from the desk setup — and this
|
||||
is the box run 3c moved to.
|
||||
|
||||
Runbook `docs/runbooks/gx10-rack-network.md`; commit `a95717e`.
|
||||
_Archived 2026-09-18._
|
||||
|
||||
- `[2026-09-03]` **Three Macs onboarded with infra-ops + NOPASSWD sudo + the DeepSeek Harness, and the fourth is a script instead of a fourth hand-run.**
|
||||
|
||||
vuongs-mac-mini 10.100.79.2 infra-ops + lkraven
|
||||
esh-macbook-air 10.0.10.83 infra-ops + lkraven
|
||||
esh-mac-studio 10.0.10.10 infra-ops + vhpfi
|
||||
|
||||
Each: key auth, `visudo`-validated NOPASSWD drop-in, password rotated to 32
|
||||
random chars and vaulted at `<name>/infra-ops-password`. `dsh` runs in the
|
||||
operator's own account on each, on a **device-scoped** LiteLLM key
|
||||
(`<name>-dsh`, scoped to `gen-reasoning`, scope verified 200/403 rather than
|
||||
trusted from the mint) — a laptop travels, and losing one should be one
|
||||
revocation, not a fleet key rotation.
|
||||
|
||||
`scripts/provision-mac-dsh.sh <host> <account> [name]` carries every trap; the
|
||||
operator-run half is `docs/runbooks/mac-provisioning.md`.
|
||||
|
||||
⚠ **`sudo -u <user>` KEEPS THE CALLER'S `$HOME`.** Without `-H` and an explicit
|
||||
`HOME=`, `"$HOME/.local"` resolved to the caller's home and an `rm -rf` aimed at
|
||||
a **working install in another account**. Only filesystem permissions stopped
|
||||
it. The script now refuses to run unless `$HOME` matches the target.
|
||||
|
||||
⚠ **An account may not own its own home.** A `sudo mkdir` before `sysadminctl`
|
||||
leaves `/Users/<account>` root-owned; the account authenticates, gets a shell,
|
||||
reports the right `$HOME`, and cannot write to it — surfacing as a bare
|
||||
"Permission denied" hours later.
|
||||
|
||||
⚠ **A wrong USERNAME looks exactly like a wrong password.** sshd answers
|
||||
`Permission denied (publickey,password,keyboard-interactive)` for a bad user, a
|
||||
bad password, AND a user outside `com.apple.access_ssh`. This produced a false
|
||||
diagnosis twice in one session — once where the password was a typo
|
||||
(`no-password` vs `nopassword`) and I blamed the access group, once where the
|
||||
Studio's operator account is **`vhpfi`, not `lkraven`**. Check
|
||||
`dscl . -list /Users` FIRST.
|
||||
|
||||
⚠ **Rotation: use `dscl . -passwd`, not `sysadminctl`.** With FileVault on and
|
||||
no Secure Token on the account, `sysadminctl -resetPasswordFor` refuses with
|
||||
"Operation is not permitted without secure token unlock". `dscl` works precisely
|
||||
because there is no token to desync. True on all three Macs.
|
||||
|
||||
⚠ **FileVault kills remote access across reboots** — the machine sits at the
|
||||
pre-boot unlock screen with no network. Nothing unattended should depend on a
|
||||
Mac being reachable after a restart.
|
||||
|
||||
⚠ macOS has no `adduser`, `useradd`, or `timeout`.
|
||||
|
||||
Harness config (all machines): `high` → the seat's `xhigh` via the gateway hook;
|
||||
`maxTokens 32768` (the 256000 default left 6144 for input and overflowed on a
|
||||
two-word prompt); `defaultContextWindow 262144`; and `models:` **replacing** the
|
||||
provider's hard-coded DeepSeek catalog, which the web GUI reads INDEPENDENTLY of
|
||||
`agent-default-model` — without it the picker offers three models the gateway
|
||||
does not serve while headless runs work fine. Commits `6ca455a`, `926fc2f`.
|
||||
_Archived 2026-09-18._
|
||||
|
||||
# `[2026-09-03]` nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding every gue
|
||||
|
||||
**nh3-dev wedged for ~40 min and it was the BACKUP, not the disk — a stalled cross-site vzdump holding every guest write via `copy-before-write`.** Symptoms screamed dying disk: 45 writes in flight completing zero, jbd2 + flush kworkers in D state 33 min, io pressure full 96%, load 26, `virtio_ring` in the stack. ⚠ **The discriminator was the ABSENCE of errors** — no SCSI/ATA/IO errors, rpool ONLINE 21%, guest fs 79%, memory fine, and **Dirty only 3.8 MB** (so nothing backed up in page cache; it was stuck BELOW the block layer). ⚠ **The hypervisor was IDLE** — load 0.63, io pressure 0.00, zpool ~0 writes: nothing was reaching the disk because the filter held it. Cause: `vzdump` of VM 102 → **pbs-ana** did 1% at 64 MiB/s then collapsed to **1.4 MiB/s for 35 min**; Proxmox backups interpose a `copy-before-write` filter, so every guest write queues behind the backup's copy-out. FIX = cancel the task (`pvesh delete /nodes/localhost/tasks/<UPID>`); filter detached, inflight 45→0, D-states gone, 191 MB/s dsync restored. ⚠ **`fleecing 0` on the job is why a slow TARGET can stall a GUEST** — fleecing routes copy-before-write to a fast local image instead. Job = `backup-5d8f1221-8f71`, **daily 21:00, `all 1`, storage pbs-ana** → recurs nightly until changed. A prior run of this VM managed 941 MiB/s read, so 1.4 MiB/s is degradation, not normal. → `docs/runbooks/nh3-dev-io-stall.md`
|
||||
_Archived 2026-09-18._
|
||||
|
||||
- `[2026-09-02]` **Every CI job on the shared `pfi-fleet` runner is root on ana-docker — and `container.valid_volumes: []` does NOT prevent it.** Measured: a job container is uid 0, `/var/run/docker.sock` is mounted by act_runner independently of that list, `docker ps` returns all 49 host containers (gitea itself, synapse, phasefinal-web, adguardhome), `docker compose v2.33.0` on PATH. ⚠ **LOAD-BEARING** — `vh/Worldtree`, `vh/soong-lab`, `vh/skaldsong`, `vh/wt-matrix-bridge` all drive buildx through that socket, so it cannot simply be closed; **isolate sensitive builds onto a dedicated runner instead.** Also measured the same night: `services:` containers work (Postgres 16), and **full-URL `uses: https://gitea.phasefinal.com/actions/checkout@v4` resolves from the local mirrors** — the un-parked half of the github-independence work, needing neither `DEFAULT_ACTIONS_URL=self` nor the act_runner auth path that blocked it on 2026-08-05. Prompted by vastblue-dev's CI-posture question for a client-funded healthcare repo. → `stacks/gitea-runner/README.md`
|
||||
_Archived 2026-09-17._
|
||||
|
||||
|
||||
Reference in New Issue
Block a user