7fdda2de53
The granite-8b harness spike proved the TRL SFT->DPO->eval seam (incl. the in-loop HoldoutEvaluator base-vs-adapter leg) runs end-to-end, but used a 12-row/12-pair synthetic writing fixture — NOT the E-RP corpus — so the ~0 anti-slop delta (-0.002) is the expected null, not an efficacy signal. Adapter reaped; nothing to A/B. Real behaviour-shift efficacy is a T1-run question. Sharpen both the T1 in-flight bullet and the Recent-decisions entry so 'green' no longer reads as efficacy-validated.
337 lines
23 KiB
Markdown
337 lines
23 KiB
Markdown
# Persistent memory — eshpfi-management
|
||
|
||
_Last updated: 2026-07-02_
|
||
|
||
## Repo purpose
|
||
|
||
Reference workspace for PFI infrastructure: server inventory, canonical
|
||
Docker Compose stacks, ops playbooks, and conventions. Authoritative
|
||
copies of compose files live on the servers under
|
||
`/opt/docker/compose/<stack>/`; this repo mirrors them for version
|
||
control, editing, planning, and CI-driven deploys. **It was originally
|
||
spun up to handle the fleet backups** — keep that lens when triaging
|
||
backup/storage issues.
|
||
|
||
## Tools and conventions
|
||
|
||
Sister repos (separate gitea repos, deployed by playbooks here):
|
||
|
||
| Repo | Role | CI status |
|
||
|---|---|---|
|
||
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
|
||
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
|
||
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
|
||
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
|
||
| `vh/althing` | Lean trusted inter-agent message bus — **v0.17 multi-machine** (re-architected 2026-06): each box runs a local-SQLite bus + a courier/receiver for P2P over the 10.x net. The v0.15 lean-bus cut RIPPED moderation / chamber (UI 7881) / forseti-daemon / agent-runner / redis-valkey. | per-box `uv tool install` (NOT CI-deploy); **nh3-extdev** added as a mesh peer 2026-06-25 (model B: althing-svc svc account + shared `/srv/althing`) |
|
||
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
|
||
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
|
||
| `vh/Worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration. **gitea-runner builds on ana-docker** (the image-bloat source); claude-bot now ADMIN collaborator (2026-06-20) | push-to-main → CI build-and-deploy (runner on ana-docker) |
|
||
| `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` |
|
||
| `vh/arbo` | Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea Actions CI (deploy-engine.sh, build-local, health-gated) now LIVE; catalog via :9009 webhook |
|
||
|
||
(`vh/volva` + Heid were re-architected from systemd daemons to Claude Code
|
||
session orchestrators 2026-06-08; their nh3-dev `.service` units were removed —
|
||
no longer deployed sidecars here. See Recent decisions.)
|
||
|
||
- **Two-layer backups** — Backrest orchestrates restic for file+DB (5
|
||
fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for
|
||
VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore +
|
||
cross-site restic targets — see `docs/runbooks/disaster-recovery.md`
|
||
for the blast-radius matrix. **⚠️ The restic file+DB layer routes
|
||
through TWO rest-servers** (`rest-server-ana` @ ana-docker:8000 →
|
||
ana-docker/ana-ml2/esh-docker-vm/vm-esh-nas; `rest-server-nh3` @
|
||
nh3-nas:8000 → irv-ml1/nh3-docker). Both depend on their NAS's NFS
|
||
export of `/mnt/backup`. (rest-server-ana recovered 2026-06-20.)
|
||
|
||
- **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace
|
||
model/dataset onto ana-ml2's shared cache at
|
||
`/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns.
|
||
|
||
- **Worldtree admin auth — per-instance.** Each Worldtree deployment
|
||
(demo :8080, personal :8081, pinned :8082) has its own Heimdall
|
||
registry and its own bootstrap admin key. Infra-ops's stored
|
||
long-lived admin key (`key_id 61419c92`) at
|
||
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`
|
||
auths against **demo only**. Personal-instance admin (the
|
||
`~/.config/worldtree/personal-admin-token`, mode 600) POSTs
|
||
`/admin/keys` (mints per-project keys; takes `user_id`+`label`, **no
|
||
scope param** — scopes are tier-derived, so admin-tier scopes like
|
||
`admin.events.read` must be granted WT-side by worldtree-dev).
|
||
|
||
- **Per-project user keys against personal Worldtree** (issued
|
||
2026-05-19): `skaldsong:79744637` (nh3-dev iteration),
|
||
`skaldsong:7c1dbbbe` (ana-docker prod), `althing:50d85460`,
|
||
`mead-hall:a360822d`. Same `user_id=skaldsong` across both
|
||
skaldsong keys → shared Heimdall agent slot; different `key_id`
|
||
→ independently rotatable. Pattern: mint via `/admin/keys`, drop
|
||
value to `/tmp/wt-personal-<name>.key` mode 600, dev collects +
|
||
shreds (DO NOT cat to chat transcript).
|
||
|
||
- **Skaldsong CD pattern (registry-pull).** Differs from althing /
|
||
asset-engine which build-on-host. vh/skaldsong's CI builds and
|
||
pushes `gitea.phasefinal.com/vh/skaldsong:<sha>` + `:latest`;
|
||
`playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates.
|
||
SHA-pin only (no `:latest` health-gated advance yet). Prereq: host
|
||
needs `docker login gitea.phasefinal.com` once (read:package PAT) —
|
||
not currently in the workflow.
|
||
|
||
- **gitea internal route for fleet hosts.** gitea is a container on
|
||
**ana-docker** — git-SSH `10.250.50.70:222`, HTTP `:3000`. Fleet/colo
|
||
hosts must use this internal route, NOT public `gitea.phasefinal.com`
|
||
(`38.120.12.44`, ana-srv1) — the public path fail2bans the host egress
|
||
IP and wedges webhook deploys. `:22` on `10.250.50.70` is ana-docker's
|
||
HOST sshd, not gitea. Full gotcha in `docs/orientation.md` → Git/gitea.
|
||
|
||
- **docker-as-root pattern** (for ops with no admin API, or to edit
|
||
deploy-owned/root-owned files without sudo): `docker run --rm -v
|
||
<target-dir>:/wt [-v /var/run/docker.sock:/var/run/docker.sock]
|
||
docker:cli sh -c "..."` (or `alpine` for plain file ops). docker-group
|
||
membership is effectively root via bind-mount; treat as sudo-equivalent.
|
||
**Foot-gun: relative paths in compose.yaml resolve against the sandbox
|
||
CWD but the Docker daemon interprets them against the HOST fs — always
|
||
pass `-e VAR=/abs/path` for any relative-default config dir.**
|
||
|
||
- **`scripts/elway` sudo handling** — elway prompts for the sudo password
|
||
ONCE via `getpass` before the first `sudo: true` step → can't run
|
||
unattended from a non-TTY tool if any step needs sudo. Sudo-free
|
||
playbooks run fully non-interactive over key SSH.
|
||
|
||
- **Per-host SSH identity matters for sudo.** infra-ops has NOPASSWD sudo
|
||
on most PFI Linux boxes (corviduo-dev included since 2026-06-15). On
|
||
ana-docker there are TWO identities: the **default `ssh ana-docker` =
|
||
`lkraven`** (docker-group, NO passwordless sudo — docker works, root-file
|
||
edits don't); **but `ssh infra-ops@ana-docker` HAS NOPASSWD root** (verified
|
||
2026-06-20). **→ For any sudo op on ana-docker (mount, root-owned files,
|
||
service control), use `ssh infra-ops@ana-docker`, NOT the default session.**
|
||
lkraven-owned files (litellm config, most stack compose/conf) still take
|
||
plain `cp`/edit under either identity. **NEW (2026-06-25):** `ssh infra-ops@10.100.10.50`
|
||
(nh3-dev) ALSO has NOPASSWD sudo; on **nh3-extdev** infra-ops is sudo-LESS by
|
||
design and **`ssh lkraven@10.100.50.42` is the NOPASSWD path** there.
|
||
|
||
## Current state / in-flight
|
||
|
||
_As of 2026-07-02:_
|
||
|
||
- **`gen` model = qwopus (`qwen3.5-122-a10b`); the Deckard trial is CONCLUDED.** Trialed
|
||
`robbatt/Qwen3.6-40B-Deckard-NVFP4` behind the `gen` aliases — it won writing quality
|
||
decisively but lost on speed (~36 vs qwopus's ~90 tok/s); reverted (git `681eb70`).
|
||
Deckard kept STAGED on ana-ml2 as T1's quality benchmark
|
||
(`/tank/aimodels/qwen36-40b-deckard-{bf16,nvfp4}`; container `vllm-deckard-40b`
|
||
stopped-but-kept). Full arc in auto-memory `reference_gen_qwopus_122b`.
|
||
|
||
- **T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS.** Attention-only LoRA (full-attn q/k/v/o
|
||
+ the bf16 GDN `in_proj_qkv`/`out_proj`); serve via swappable-LoRA-on-NVFP4 (test QUEUED,
|
||
gated on the first T1 adapter) or merge+requant. Trainer harness seam MECHANICALLY
|
||
proven via a granite-8b spike (2026-07-02): the pipeline runs end-to-end, but efficacy
|
||
is NOT validated (mechanics-only synthetic fixture, by design; a granite adapter
|
||
wouldn't transfer to qwopus regardless). Real behaviour-shift efficacy is a T1-run
|
||
question. NVFP4 quant-structure + serve-path facts in `reference_gen_qwopus_122b`.
|
||
|
||
- **Worldtree #332 scoped-log view + tunnel = STANDING ASSET.** `wt_gateway_logs` view +
|
||
`wt_readonly` role on the litellm DB + `wt-db-tunnel` systemd on corviduo-dev, for
|
||
worldtree-dev's embed-recall regression-watches. #332 fix verified in prod (15×→1.01×
|
||
re-embed). Teardown steps + the IP-pin caveat in auto-memory
|
||
`reference_wt_gateway_scoped_log_view`.
|
||
|
||
- **`/books` mounted (transient) on nh3-dev** for a books/corpus ingestion:
|
||
`10.0.50.50:/mnt/books` (ESH NAS) → `/mnt/books`, NFSv4 `ro,soft` (soft dodges the ESH
|
||
D-state hang). NOT fstab — re-mount via `infra-ops@10.100.10.50` if nh3-dev reboots.
|
||
|
||
- **granite-4.1-8b bf16 kept** at `irv-ml1:/home/lkraven/granite-4.1-8b-bf16` (17G,
|
||
reap-on-request; the T1-harness-spike base). comfyui was borrowed off the A6000 for the
|
||
spike + RESTORED healthy.
|
||
|
||
- **Recently completed (2026-06-22..25, now in Recent decisions):** Worldtree #314/#322/#317
|
||
persona-render config arc; nh3-extdev althing v0.17.1 mesh peer (Model B); zellij web-seat pilot.
|
||
|
||
- **Backups — recovered + hardened (2026-06-20), STILL OPEN:** rotate the 5 disclosed
|
||
rest-server creds (operator, offline); confirm esh-vm-db's resticprofile includes DB
|
||
dumps (esh-pve-nas restic-only). Topology + 2-min check: `docs/runbooks/backups.md`.
|
||
|
||
- **R22 (brokkr/dwarves) — CONCLUDED, gateway-only.** Phase B (internal-model hooks)
|
||
CANCELLED (Worldtree model-agnostic → no deploy path). Whole footprint = gateway calls
|
||
to free `qwen3.5-122-a10b` via a PERSISTENT full-access key at
|
||
`/home/lkraven/.r22-gateway-key` (mode 600, carries paid GLM, do NOT delete). Detail in
|
||
Recent decisions.
|
||
|
||
- **Open follow-ups (low-priority):** clean phantom `qwen3.6-35b-a3b` off the gateway
|
||
(lists in /v1/models but 400s — backend displaced); docker-daemon default log-cap on
|
||
ana-docker (systemic fix behind the 94 GB disk incident).
|
||
|
||
- **Standing / parked (from prior):** Mac Pro migration (`migration-plan.md`,
|
||
hardware-gated); R17 v2 corpus push HELD (distribution barred); disclosed-keys rotation
|
||
queue; clean legacy `news-digest` on ana-docker.
|
||
|
||
- **Worldtree config-propagation (reference):** demo+personal bind-mount config
|
||
(`providers.yaml`/`model_roles.yaml`/`policies.yaml`/`defaults.yaml`, now incl. the
|
||
#314/#322/#317 additions) from `/opt/worldtree{,-personal}/config`
|
||
(infra-ops-deployable, byte-identical from worldtree-dev canonical sha); pinned is
|
||
IN-IMAGE. Reload via `docker restart <container>`, NEVER `compose up` (stale-`:latest`
|
||
footgun); WT PDP is rule-based (a tier needs allow RULES, not just scopes). **#317
|
||
lesson: config REMOVALS are NOT backward-compatible with the still-running image — push
|
||
promptly, don't bounce the instance in the window.**
|
||
|
||
## Recent decisions
|
||
|
||
- `[2026-07-02]` **mtf-dev granite harness-spike provisioned on irv-ml1 (comfyui displaced
|
||
for the A6000) → ran GREEN.** Freed the A6000 by stopping comfyui (operator-coordinated
|
||
with comfy-dev), staged granite-4.1-8b bf16 to `irv-ml1:/home/lkraven`, mtf-dev's TRL
|
||
SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the
|
||
expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's
|
||
trainer harness ahead of the real qwopus train. **MECHANICAL green ONLY — efficacy NOT
|
||
validated, by design** (mtf-dev confirm 2026-07-02): the run used a 12-row/12-pair
|
||
SYNTHETIC writing fixture (clean-vs-sloppy generic prompts), NOT the E-RP corpus; rank 8,
|
||
1 epoch, ~2 SFT + ~6 DPO steps. The in-loop base-vs-adapter check (HoldoutEvaluator)
|
||
already ran and returned the expected null (anti_slop_improvement −0.002); the adapter was
|
||
then reaped (gone), so there is nothing to A/B — and on a mechanics-only synthetic adapter
|
||
an A/B would only reconfirm the null. Real efficacy = the T1 run (real recipe + E-RP data
|
||
on qwopus). Open (mtf-dev routing to operator): whether to insert an intermediate
|
||
*real-efficacy* granite spike before T1 — weak proxy (granite arch ≠ qwopus, + granite is
|
||
censored vs qwopus abliterated), leaning defer-to-T1.
|
||
|
||
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel
|
||
provisioned + fix verified.** The persona-recitation gate re-embedded persona segments
|
||
every turn (15× re-embed, 95% cross-turn recurrence, one anchor ×206/min → worldtree-gateway
|
||
was ~95% of the embedding backend load); worldtree-dev's cross-turn content-hash cache
|
||
dropped it to 1.01× in prod. Built a boundary-safe read-only `wt_gateway_logs` view +
|
||
`wt_readonly` role + `wt-db-tunnel` (corviduo-dev) for their regression-watch. auto-memory
|
||
`reference_wt_gateway_scoped_log_view`.
|
||
|
||
- `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept.** qwopus
|
||
HAS a full native MTP head (785 tensors, in NVFP4 + bf16). Measured +12% single-stream but
|
||
**−15–20% AGGREGATE at moderate concurrency** (N=4: 250→200 tok/s) + it silently drops
|
||
`min_p`/`logit_bias` → reverted to clean baseline. The win is banked for T1 (MTP as a
|
||
per-deployment option, preserved through requant). `reference_gen_qwopus_122b`.
|
||
|
||
- `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`).** `robbatt/Qwen3.6-40B-Deckard-NVFP4`
|
||
won writing "in every way" but at ~36 vs ~90 tok/s (dense-40B vs MoE-~10B-active);
|
||
spec-decode rescue ruled out (aeon DFlash image is arm64/DGX-Spark-only; Deckard's
|
||
NVFP4/bf16/base-27B all LACK MTP; EAGLE-head training a multi-week non-starter). git
|
||
`b63c48b` (repoint) + `681eb70` (revert). Deckard kept staged as T1's writing benchmark.
|
||
|
||
- `[2026-06-25]` **althing re-architected to the v0.17 lean multi-machine bus; nh3-extdev
|
||
stood up as a mesh peer (MODEL B).** v0.15.0's lean-bus cut RIPPED
|
||
moderation/chamber/forseti-daemon/agent-runner/redis-valkey; **v0.17 = per-box
|
||
local-SQLite bus + a courier/receiver for P2P over the 10.x net** (installed per-box via
|
||
`uv tool install`, NOT CI). Migrated nh3-dev's client v0.15→v0.17. nh3-extdev
|
||
provisioned as a full mesh peer with the operator's **MODEL B** (dedicated `althing-svc`
|
||
service account + group-shared `/srv/althing` root, so multiple OS users share one
|
||
config/DB; `ALTHING_ROOT` via /etc/profile.d). Verified multi-user concurrent rw.
|
||
Supersedes the old althing Tools-row (was chamber/daemons/valkey). auto-memory
|
||
`reference_nh3_extdev_althing_mesh`.
|
||
|
||
- `[2026-06-23]` **zellij native web client piloted on nh3-dev** (`zellij-web.service`
|
||
:8443, native TLS), system-installed for all users, ALONGSIDE ttyd (kept as the no-auth
|
||
fallback). Bind left `0.0.0.0` per operator. auto-memory `reference_zellij_web_seat`.
|
||
|
||
- `[2026-06-22]` **Worldtree persona-render config arc (#314 relational-stance / #322
|
||
affect-mood gate / #317 valence teardown) pre-synced + deployed green on demo+personal**
|
||
— #317 a boot-blocking config REMOVAL (relational_valence). Config-sync recipe +
|
||
additions-vs-removals lesson in auto-memory `reference_corviduo_dev_emergency_ops`.
|
||
|
||
- `[2026-06-20]` **R22 (brokkr/eitri/dwarves) stood down to gateway; full-access R22
|
||
key minted; Phase B parked.** Operator caught the raw-weights/GPU ask as premature —
|
||
Phase A (P00 + NL-rendering) is PROMPT-LEVEL on the LiteLLM gateway, not raw weights
|
||
(no GPU, no Blackwell displacement). Real blocker was gateway ACCESS: the shared
|
||
`all-agents-local` key only serves granite chat, and granite-8b is too weak to be the
|
||
P00 model-under-test (capability-confound). Operator said "full access" → minted
|
||
`r22-brokkr-phaseA` LiteLLM key (scope **all-proxy-models**, incl PAID GLM), dropped
|
||
`~/.r22-gateway-key` (mode 600) for eitri — **re-minted PERSISTENT** after eitri's
|
||
stateless tool-env couldn't hold it in env + shredded the one-shot drop (old key
|
||
revoked); the live full-access key now lives at `/home/lkraven/.r22-gateway-key`,
|
||
read per-command, **don't delete** (cost surface — carries paid GLM; operator
|
||
reconfirmed full-open + declined to narrow, 2026-06-20 — settled, don't re-surface).
|
||
Lesson: stateless consumers need a persistent mode-600 file read per-call, NOT
|
||
shred-after-load. **P00 MUT = `qwen3.5-122-a10b` (alias
|
||
`gen`)** — 122B, capable, FREE-LOCAL (no per-token cost); told them keep frozen
|
||
calibration on it, NOT the paid GLM the key can now reach. **PHASE B** (raw-weights
|
||
hooks: Dvalin activation-steering on hidden states + Regin ECM logit-injection on a
|
||
7-8B Qwen/Llama via transformers LogitsProcessor) **CANCELLED 2026-06-20 (not
|
||
deferred)** — operator cut the internals arms: **Worldtree is model-agnostic** (regard/
|
||
affect renders via context manipulation = prompt-level only), so the internal-model
|
||
approach has no deploy path (the pragmatic-outcome gate firing). R22's ENTIRE compute
|
||
is now the gateway (free qwen3.5-122-a10b MUT); NO raw-weights/GPU ever; the A6000
|
||
slice reservation is released (was never provisioned). The full-scope R22 key stays
|
||
(pinned to the free qwen-122, GLM untouched). **OPERATOR STEER (2026-06-20): R22 is
|
||
research for a PRAGMATIC/deployable outcome, NOT advancing-the-art.** Before the dwarves
|
||
invest in Phase B, establish a deploy path (distilled steering-vector / fine-tune
|
||
applied to a model we already serve). Operator is taking the dwarf discussion directly.
|
||
|
||
- `[2026-06-20]` **claude-bot issue-scope token minted for worldtree-dev self-serve**
|
||
(closes their last `tea`-as-vh fallback — issues; they already self-serve Actions +
|
||
deploys). Gitea PAT scopes are IMMUTABLE → minted a NEW claude-bot PAT
|
||
`worldtree-dev-actions-issues-20260621` (id 16) with `write:repository`+`write:issue`
|
||
via basic-auth (`claude-bot:gitea-password`, POST `/users/claude-bot/tokens`); dropped
|
||
mode-600 `~/.claude-bot-token-issues` → worldtree-dev swaps. **Gitea COLLAPSES
|
||
`read:issue` into `write:issue`** (write implies read). Old token (id 15) REVOKED after
|
||
verification → claude-bot carries ONE consolidated worldtree-dev token (id 16) + arbo-ci
|
||
(id 14). Advances the credential-migration directive.
|
||
|
||
- `[2026-06-20]` **rest-server-ana recovered + backup prevention shipped + worldtree-dev
|
||
admin keys provisioned** (operator-directed). rest-server-ana fixed (mount-a +
|
||
force-recreate via `infra-ops@ana-docker`); freshness alert + fstab hardening landed
|
||
(see `docs/runbooks/backups.md`). Minted worldtree-dev **admin-tier Heimdall keys** on
|
||
demo (key_id d113207c) + personal (f4f75adb), dropped mode-600 to
|
||
`~/.wt-admin-{demo,personal}` on nh3-dev → worldtree-dev now self-serves key minting.
|
||
Cred rotation (5 rest-server pw) BELAYED per operator.
|
||
|
||
- `[2026-06-20]` **claude-bot → ADMIN on vh/Worldtree** (operator-authorized;
|
||
one-time use of operator vh-admin to enable the migration) — claude-bot
|
||
self-serves WT deploys/tokens henceforth, reducing personal-cred fallbacks.
|
||
|
||
- `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.** Stand up service accounts (claude-bot Gitea user + scoped tokens, distinct litellm admin key); re-mint consumer creds under them; flag personal-cred fallbacks until done. (auto-memory `project_migrate_infra_access_to_claude_credentials`)
|
||
|
||
_118 older entries archived to archival-memory.md._
|
||
|
||
## Tried and abandoned
|
||
|
||
- `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found /
|
||
unauthorized" is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml`
|
||
is PULL-ONLY ("image must already be built by a push to main") but fires on the
|
||
`staging/vX` tag push SIMULTANEOUSLY with `deploy.yml`'s main build → it tries to pull the
|
||
image ~2.5 min BEFORE the build finishes pushing it → step-4 "Verify image exists" aborts
|
||
"not found". Misattributed to registry-auth twice (the earlier cff3328 saga too). DIAGNOSE:
|
||
the image tag (12-char short-sha, NOT 7) exists in the registry + the VM login succeeds ⇒
|
||
it's the race. FIX: re-run once the build's done (image now present), OR gate deploy-personal
|
||
on `workflow_run: completed`.
|
||
|
||
- `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS
|
||
moderate-concurrency aggregate throughput + silently ignores `min_p`/`logit_bias`.**
|
||
Measured on qwopus `gen`: N=1 +12%, N=4 −20% aggregate. Don't bolt spec-decode onto the
|
||
fleet `gen` for a single-stream win — reserve it for dedicated/interactive deployments.
|
||
|
||
- `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (and
|
||
irv-ml1 sudo needs a password non-interactively) → stage model pulls to `/home`.** Also
|
||
PIN THE A6000 BY UUID for training runs: nvidia-smi index 1 is the A6000, but native-CUDA
|
||
ordering can differ vs docker, and the 3090 (index 0) is usually near-full → land there and
|
||
OOM. `CUDA_VISIBLE_DEVICES=GPU-<uuid>`.
|
||
|
||
- `[2026-06-25]` **althing "unreachable: <machine> — retry later" can MASK an app-level
|
||
500.** A multi-hour "intermittent :8087 / firewall flap" hunt was a red herring: raw
|
||
network was always clean (curl POST to the receiver `:8087` worked; a connect probe = 0
|
||
fails). Root cause = the receiver's DB agents-table wasn't synced with the config roster,
|
||
so delivery 500'd `"unknown to: <handle>"`, which the courier MAPPED to "unreachable"
|
||
(looks like network/DNS). **Diagnose: raw curl to :8087 + a connect-probe pass ⇒ NOT
|
||
network — look up-stack.** Fixed in althing **v0.17.1** (receiver auto-ensures the
|
||
recipient on delivery). Stop-gap on older builds: `ALTHING_HANDLE=<h> althing-cli inbox`
|
||
as the receiver user syncs config→DB; re-run after any roster change. auto-memory
|
||
`reference_nh3_extdev_althing_mesh`.
|
||
|
||
- `[2026-06-20]` **rest-server `.htpasswd: permission denied` = the ana-nas
|
||
NFS mount FAILED (ghost file on the local mount point), NOT a decommission.**
|
||
`mnt-backup.mount` stuck `failed` (fstab bare `defaults`, no retry) → the
|
||
rest-server serves an empty local dir with a root:root 0-byte `.htpasswd`.
|
||
Documented recovery in disaster-recovery.md. Don't tear down what looks like
|
||
a crash-looping legacy container until you've checked fstab + the docs — it
|
||
was the live ana-side restic target.
|
||
|
||
- `[2026-06-20]` **The DEFAULT `ssh ana-docker` is `lkraven` (no NOPASSWD) —
|
||
but `ssh infra-ops@ana-docker` HAS NOPASSWD root** (corrected later same day).
|
||
Early on a `sudo cp` as lkraven silently failed (password prompt) → one
|
||
wasted gateway bounce; I then wrongly concluded "no sudo on ana-docker" and
|
||
nearly punted the rest-server-ana recovery to the operator. The real rule:
|
||
reach for `infra-ops@ana-docker` for sudo ops; lkraven-owned files (litellm
|
||
config, most stack compose/conf) take plain `cp` under either identity.
|
||
|
||
_98 older entries archived to archival-memory.md._
|