Files
esh-pfi-infrastructure/persistent-memory.md
T

380 lines
27 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Persistent memory — eshpfi-management
_Last updated: 2026-06-13_
## Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
`/opt/docker/compose/<stack>/`; this repo mirrors them for version
control, editing, planning, and CI-driven deploys.
## Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
| `vh/althing` | Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) | push-to-main → CI deploys (2026-05-14) |
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
| `vh/worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration | push-to-main → CI deploys |
| `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` |
(`vh/volva` + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev `.service` units were removed —
no longer deployed sidecars here. See Recent decisions.)
- **Two-layer backups** — Backrest orchestrates restic for file+DB (5
fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for
VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore +
cross-site restic targets — see `docs/runbooks/disaster-recovery.md`
for the blast-radius matrix.
- **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace
model/dataset onto ana-ml2's shared cache at
`/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns.
- **Worldtree admin auth — per-instance.** Each Worldtree deployment
(demo :8080, personal :8081, pinned :8082) has its own Heimdall
registry and its own bootstrap admin key. Infra-ops's stored
long-lived admin key (`key_id 61419c92`) at
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`
auths against **demo only**. For personal-instance admin ops, fetch
the bootstrap admin per-op via
`docker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEY`
on corviduo-dev. Used for `POST /admin/keys`, admin diagnostics
(`/admin/sessions/<id>/{bifrost,tools}`, etc.).
- **Per-project user keys against personal Worldtree** (issued
2026-05-19): `skaldsong:79744637` (nh3-dev iteration),
`skaldsong:7c1dbbbe` (ana-docker prod), `althing:50d85460`,
`mead-hall:a360822d`. Same `user_id=skaldsong` across both
skaldsong keys → shared Heimdall agent slot; different `key_id`
→ independently rotatable. Pattern: mint via `/admin/keys`, drop
value to `/tmp/wt-personal-<name>.key` mode 600, dev collects +
shreds (DO NOT cat to chat transcript).
- **Skaldsong CD pattern (registry-pull).** Differs from althing /
asset-engine which build-on-host. vh/skaldsong's CI builds and
pushes `gitea.phasefinal.com/vh/skaldsong:<sha>` + `:latest`;
`playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates.
SHA-pin only (no `:latest` health-gated advance yet). Prereq: host
needs `docker login gitea.phasefinal.com` once (read:package PAT) —
not currently in the workflow.
- **docker-as-root pattern** (for ops that have no admin API, e.g.
`SqliteUserStore.set_bifrost_credentials`): on hosts where the SSH
user is in the `docker` group but lacks passwordless sudo, run
`docker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..."` to edit deploy-owned files
without sudo. Documented with security warning in
`servers/corviduo-dev/README.md`. docker-group membership is
effectively root via bind-mount; treat as a sudo-equivalent grant.
**Foot-gun: when running `docker compose` inside this sandbox,
any relative path in compose.yaml (e.g. `${WORLDTREE_CONFIG_DIR:-./config}`)
resolves against the sandbox CWD, but Docker daemon interprets the
resulting path against the HOST filesystem. Always pass `-e VAR=/abs/path`
to the docker run invocation for any relative-default config dir.**
- **`scripts/elway` sudo handling** — elway prompts for the sudo password
ONCE via `getpass` before the first `sudo: true` step. That prompt is
interactive → elway can't run unattended from a non-TTY tool if any step
needs sudo. For sudo-free playbooks (no `sudo: true` steps) it runs fully
non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo,
use the docker-daemon-root trick: `docker run --rm -v /worktank:/mnt alpine
sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'`.
## Current state / in-flight
_As of 2026-06-13:_
- **ana-ml2 went Ada → Blackwell** (dual NVIDIA RTX PRO 6000 Blackwell Max-Q, 96 GB each, cc 12.0 /
sm_120 — was dual RTX 6000 Ada 48 GB / cc 8.9, confirmed live via `nvidia-smi`). GPU layout now:
**GPU 0** held free for large-model hot-loads (llama-swap pinned there, `edf0f91`); **GPU 1** is the
steady-tenant card — granite (131k ctx) + qwen3.5-vl (65k) + the embed/rerank/reward trio, ~3.5 GB
free after the rebalance. CLAUDE.md's server table + the granite/qwen "Ada (cc 8.9)" compose comments
are STALE — doc-fix offered, **pending operator go-ahead** (/snapshot doesn't auto-edit CLAUDE.md).
- **R16 "vmoan" TTS LoRA POC — awaiting the operator's ear** (Brokkr's designated gate). Two auditions
on NFS: `/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v1.wav` (42 s, run-on) vs `_v2.wav` (10 s,
tight). v2 fixed the run-on; Brokkr holds the pilot verdict on the operator's listen. Grab:
`scp irv-ml1:/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v{1,2}.wav .`. LOCAL-ONLY corpus
(soundgasm-derived) — distribution barred; never echo transcripts to the bus.
- **R17 v2 corpus characterization still on the local irv-ml1 branch, push HELD** (`r17-v2-character
ization`, 18 MULTI / 101 ELIGIBLE + 19936 candidates + `speech.json` ASR). No-push rule AND ASR of
the intimate-audio batch (content-exposure call). Awaiting brokkr-collect or explicit push approval.
- **Mac Pro migration is still the big open project** — `migration-plan.md` (repo root): WORKSTATION-
ONLY move of nh3-dev's Hat-1 dev env (~40 repos, all of `~/.claude`, dotfiles, toolchain) to an M2
Ultra Mac Pro Rack on the 10.100 subnet. Hat-2 fleet sidecars (egress SOCKS5, mead-hall, ttyd seats,
`althing-forseti`) STAY on the Linux VM. **Phase 0** (push pushables + confirm the ~6 local-only repos)
is the only time-sensitive step. Cutover = `rsync` working trees, NOT re-clone (re-clone loses unsaved
work in ~15 repos). macOS `/home/lkraven`→`/Users/lkraven` repath. Everything else waits on hardware.
- **sglang-vs-vLLM bench stack staged but parked** (`5f049cb`) — the originating question (SGLang
RadixAttention caching) resolved without it: vLLM v1 already defaults prefix-caching ON, now pinned
explicit on granite+qwen. The NVFP4 chase is abandoned. Bench stack stays staged if ever revisited;
the MAX_JOBS-on-shared-prod foot-gun (see Tried) caps any future from-source build on ana-ml2.
- **pi + GLM 5.1 harness live on nh3-dev** — `glm` launcher runs pi against `glm-5.1` (thinking-off)
via the LiteLLM gateway; `glm-5.1-reasoning` for opt-in thinking.
- **Worldtree config-propagation lane is mature + humming** — pre-merge delta-ping → sync-on-merge to
the demo+personal bind-mounts; pinned stays pre-cutover (full migration at its next re-image).
- **Disclosed-keys hygiene queue** (rotate at convenience): HF token, the wt-personal keys, Gitea
runner reg token, MINIFLUX_PASSWORD, ZAI_API_KEY. (The shared all-agents LiteLLM key is intentional,
not hygiene-debt — see decisions; rotatable via infra-ops only if it leaks.)
- **Still open from prior:** clean legacy `news-digest` on ana-docker; the `docker push 60s ceiling`
mystery uninstrumented. (llama-swap GPU-0 pin — DONE this session, `edf0f91`.)
## Recent decisions
- `[2026-06-13]` **ana-ml2 upgraded Ada → dual RTX PRO 6000 Blackwell Max-Q** (96 GB each, cc 12.0 /
sm_120; was dual RTX 6000 Ada 48 GB / cc 8.9 — confirmed live via `nvidia-smi`). Unlocks NVFP4 (FP4
tensor cores) and doubles VRAM headroom. CLAUDE.md's server table + the granite/qwen compose
"Ada (cc 8.9)" comments are now stale; doc-fix **offered, pending operator go-ahead** (snapshot doesn't
auto-edit CLAUDE.md). Tracking surface: this snapshot + commits `19a07b9`/`1e2a3a1` ("Blackwell 96GB").
- `[2026-06-13]` **NVFP4-W4A4 is infeasible for Granite — FP8 stays the Granite-on-Blackwell format.**
W4A4 (4-bit weights + activations) collapses at 30k context, proven **producer-independent** (modelopt
AND llm-compressor both clean-NONE from the same BF16 base + wikitext-2k calib). No 4-bit wins both
axes: W4A4 = quality collapse; W4A16-NVFP4/AWQ = weight-only dequant → bf16 (no FP4-core speedup), so
the tempting ~40% uplift doesn't materialize without quality loss. **30B retired** (not enough quality
for the VRAM/perf hit for our use case). (auto-memory `reference_nvfp4_w4a4_granite_infeasible`)
- `[2026-06-13]` **Qwen3.5-9B VL (FP8) deployed on ana-ml2 GPU 1** — `qwen35-vl` stack, :8007, gateway
alias `qwen3.5-9b-fp8`. Clean FP8 (no NVFP4 for vision). **Pinned nightly digest, not `:latest`**: the
stable release quantizes the Qwen3.5-VL *vision tower* under `--quantization fp8` → garbage vision (LM
fine, "sees" noise); the nightly correctly excludes the vision tower. Re-pin to `:latest` + drop the
pin once that exclusion lands stable. (`2e3dcc2`)
- `[2026-06-13]` **comfyui 325 G model tree migrated worktank → `/storetank/arbo`** (worktank 97% → 26%).
`arbo` is the consuming app (ComfyUI-backed prompt/image pipeline); models live under its name on the
bigger pool. Overlay bind-mount via `COMFYUI_MODELS_DIR` in the comfyui compose; inventory at
`docs/arbo-comfyui-model-catalog.md` (retain-large-portion-for-future-use decision). (`38186be`)
- `[2026-06-13]` **GPU layout settled on the Blackwell box.** GPU 0 held free for large-model hot-loads
(llama-swap pinned there, `edf0f91`); GPU 1 is the steady-tenant card — granite 131k ctx (was 51k),
qwen 65k, embed/rerank/reward trio, ~3.5 GB free after rebalance (`1e2a3a1`, `19a07b9`; trio re-floored
for 96 GB, 20×-parallel-stable). embed/rerank left at floor — long docs are chunked BEFORE embedding,
so a longer embed ctx buys nothing. PagedAttention note: max-model-len is a ceiling, not a reservation,
so small requests aren't blocked by the big ceiling; concurrency = KV-pool-tokens / actual-request-size.
- `[2026-06-13]` **Prefix caching pinned explicit on granite + qwen** — benched ~6.5× faster TTFT (45 ms
vs 292 ms) on a shared ~4.5k-token summarizer template; soft/evictable KV, neutral when prefixes don't
repeat. vLLM v1 `:latest` defaults it ON (granite) but the qwen nightly defaults it OFF — pin both so a
version flip can't silently disable it. (`a9a2be7`)
- `[2026-06-13]` **granite-4.1-8b listed as the always-available summarizer/classifier endpoint + a
shared all-agents key minted** (operator-directed). Added to the global `~/.claude/CLAUDE.md` Global-
tools section; key alias `all-agents-local`, scoped to the FREE local models only (granite + qwen-
vision + embed/rerank, NOT the paid GLM), internal-gateway-only, rotatable. The standing "reach for
this before spending premium tokens on low-caliber high-volume work" lever. (auto-memory
`reference_litellm_gateway`)
- `[2026-06-13]` **arbo engine + frontend stack stood up** (ADR-0001) — irv-ml1 co-located inference
engine (`ee57e69`), python-based healthcheck (the slim image ships no curl/wget, `bdb3312`), frontend
ro-mounted from the v0.11.2 checkout (`922e8ad`, ADR-0001 D2). arbo = the ComfyUI-consuming app whose
models now live at `/storetank/arbo`.
- `[2026-06-11]` **GLM thinking inverted at the LiteLLM gateway** (operator call): `glm-5.1` defaults
thinking-OFF; new `glm-5.1-reasoning` alias = same z.ai upstream with thinking ON (opt-in).
Mechanism: `litellm_params.extra_body:{thinking:{type:disabled}}` — LiteLLM `drop_params` STRIPS a
top-level `thinking`/`reasoning_effort`, but forwards `extra_body` verbatim to z.ai (the only channel
that works; verified reasoning_tokens 0 vs >0). Shared-gateway change — affects ALL glm-5.1 callers
(brokkr's all-models key included). (`95b2701`, auto-memory `reference_litellm_gateway`)
- `[2026-06-11]` **pi coding agent installed on nh3-dev as a GLM 5.1 harness** — `@earendil-works/
pi-coding-agent` via **bun** (user-level; npm's global prefix is `/usr` → needs sudo, bun avoids it).
Config `~/.pi/agent/models.json` (litellm provider → gateway), launcher `~/.local/bin/glm` sources
the gateway key + selects the model. pi is OpenAI-compatible; proxy-safe compat flags for the GLM path.
- `[2026-06-11]` **z.ai web-tools (regin) = z.ai hosted MCP path, NOT the `/paas/v4` Tool API.** Direct
`api.z.ai/api/paas/v4/web_search` → 429/1113 "insufficient balance" (coding-plan keys bill tools on a
separate quota path); `open.bigmodel.cn` is the China platform (404, different account). WORKS: MCP
streamable-HTTP at `https://api.z.ai/api/mcp/{web_search_prime,web_reader}/mcp`, `Authorization: Bearer
$ZAI_API_KEY` (the **MCP** key — distinct from `Z_AI_API_KEY` the LLM key). Reference impl = Worldtree's
Leif agent (`tools/web/zai_client.py` + `core/clients/mcp.py`).
- `[2026-06-10]` **Mac Pro migration framed: workstation-only** (M2 Ultra ARM, racked NH3 on-subnet);
sidecars stay Linux. `migration-plan.md`. (See in-flight.)
- `[2026-06-10]` **Worldtree deployed-config propagation is infra-ops's OWNED lane** (operator ruling).
worldtree-dev pings the config delta on every config-touching commit (pre-merge); infra-ops syncs
`config/*.yaml` from MERGED canonical to the `/opt/worldtree*/config` bind-mounts on demo+personal
(Heimdall hot-reloads policies; defaults are code-defaulted). The **v0.33.8 9-HOUR demo outage** — a
`model_roles.yaml` hard-startup-dep that shipped in canonical 3 releases earlier but never reached the
VM — is the failure mode this lane prevents. providers.yaml stays hand-tuned (artemis graft on
personal). corviduo-dev emergency-ops = `ssh vh@10.250.50.152` (alias doesn't resolve, lkraven denied,
infra-ops key excluded), docker no-sudo. (auto-memory `reference_worldtree_deploys_cicd`,
`reference_corviduo_dev_emergency_ops`)
- `[2026-06-09]` **LiteLLM scoped virtual keys issued to consumers** (operator-authorized): `brokkr-
smithy` (all-proxy-models), `arbo-prompt-enhance` (comfy-dev, granite-only). Pattern: mint via
`/key/generate` (master `sk-corvid`), scope-restricted + rotatable, value → 600 file never the bus.
(auto-memory `reference_litellm_gateway`)
- `[2026-06-08]` **volva.service + heid.service removed from nh3-dev** — vestigial systemd daemons;
Heid/Volva re-architected from Python pollers to Claude Code session orchestrators (heid commit
`12aa5a9`); binaries+dirs gone, volva.service was crash-looping 203/EXEC. Cleanup at heid's request
(the bus-content-driven sudo got the harness guardrail; operator green-lit). (`6e2f80e`)
- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer** (supersedes
the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not
Q4/AWQ** — primary workload (agent memory + summarization) is high-concurrency, where FP8 scales
~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a config bug).
vLLM `vllm-granite` :8004 GPU 1, official IBM compressed-tensors FP8, CUDA graphs. (Then on Ada cc 8.9;
the box has since gone Blackwell — see 2026-06-13.) (`34a43a0`, auto-memory `reference_ana_ml2_vllm_granite`)
- `[2026-06-05]` **Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI**; LiteLLM
`success_callback:[langfuse]` live (project `gateway`). Pretty prompt/completion/reasoning traces +
an `outputTokensPerSecond` tok/s dashboard. NOT a prerequisite — spend_logs already capture
tokens+latency. (`9171e6a`, auto-memory `reference_litellm_gateway`)
- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any
found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memory
`feedback_avoid_ollama`)
- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent).
FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI;
eshpfi keeps the `comfyui` stack compose, comfy-dev owns the model/workflow knowledge. (auto-memory
`reference_irv_ml1_ampere_quant`)
- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is
the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to
claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands
the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts.
**CORRECTION to the 2026-06-04 "deploys ALL CICD" line:** the bind-mount CONFIGS (providers.yaml,
vh-owned on corviduo `/opt/worldtree*/config`) ARE infra-ops's to apply directly — only the
app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`)
- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent;
granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from
Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`)
- `[2026-06-04]` **phi4 ships the CANONICAL/official Phi-4 chat template, NOT Ollama's.**
Ollama's bundled template omits the system `<|end|>` — that flattered brokkr's R15 eval but is
the DIVERGENT scaffold (Dvalin: the system `<|end|>` is Microsoft's intended format). Applied an
Ollama-matching override then reverted — ship correct, not the benchmark quirk. (`90e08f0`→`27eb537`;
"headgun" lesson in Tried.)
- `[2026-06-04]` **infra-ops NOPASSWD-sudo identity commissioned, scoped to PFI boxes** (+esh-docker-vm
by operator override) — so infra-ops completes DevOps end-to-end vs handing the operator sudo steps.
Dedicated key, sudo log_output, key-gated. (`8c32a05`)
- `[2026-06-04]` **Worldtree demo/pinned/personal deploys are ALL CI/CD, not infra-ops** — a "deploy
vX.Y.Z" request to infra-ops is MISROUTED → point them back to their pipeline. The granite→phi4
repoint: worldtree-dev self-served via their CI/CD (v0.30.10). (`d8d776c`)
- `[2026-06-04]` **ollama upgraded 0.9.0→0.30.4 on irv-ml1** (Ministral-3 is a Dec-2025 model the
old engine refused); A6000 pinned by **UUID** not index (native fastest-first ≠ nvidia-smi PCI).
- `[2026-06-04]` **`brokkr` user (no-sudo) on irv-ml1; R14/R15/R16 substrate moved to /home/brokkr.**
Persistent box services there need SYSTEM systemd units (see Tried).
_44 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-06-13]` **Heavy from-source compile (`MAX_JOBS=128` sglang fork build) on the shared PROD GPU
box PINS it** — load hit 187, prod vLLM services restarted, killed an in-flight quant. ana-ml2 hosts
live inference; never run a big build there at full parallelism. Cap `MAX_JOBS≤32`, build off-box, or
cgroup-constrain. (Operator ran the kill; infra-ops NOPASSWD-sudo confirmed working on ana-ml2 — retry
infra-ops on an ssh-255 before concluding "no access.")
- `[2026-06-13]` **`--quantization fp8` on a VL model can quantize the VISION TOWER → garbage vision**
(Qwen3.5-VL on the stable vLLM: gray-grid output; the LM answers text fine, so it "looks" healthy
until you actually feed it an image). The nightly excludes the vision tower. Lesson: validate the
VISION path on a quantized VLM, not just text — and pin the engine digest that has the exclusion.
- `[2026-06-13]` **vLLM's `--gpu-memory-utilization` is checked against FREE VRAM at startup
(`free >= util*total`), not total** — on a shared card, growing one service before trimming a
co-tenant OOMs ("free 48.56 < desired 80.72" at util 0.85 on a half-occupied 96 GB card). Start-order
matters: trim the shrinking service FIRST, then grow the other. Size to the FREE budget, not the total.
- `[2026-06-13]` **The `vllm/vllm-openai` entrypoint is already `["vllm","serve"]`** — the compose
`command:` supplies the model as the first POSITIONAL arg + flags; a second `serve` (or `--model X`)
yields "unrecognized arguments". Same-class gotchas this session: `tee` masks the real exit code (use
a `>` redirect to keep rc); HF `datasets` rejects bare `wikitext` (needs `Salesforce/wikitext`).
- `[2026-06-13]` **Chatterbox-Turbo LoRA finetune: the repo's `setup.py` loads the WRONG tokenizer** —
`merge_and_save_turbo_tokenizer()` pulls gpt2-medium + a grapheme merge file (len mismatch) instead of
the chatterbox-turbo GPT2 tokenizer (vocab.json+merges.txt, len 50276). Fix = override with the correct
tokenizer + delete the grapheme `tokenizer.json`; `[vmoan]` token → new_vocab_size 50277 (1-row resize),
lora_r 64 / alpha 128, modules_to_save=[text_emb,text_head]. Also: a unique-stem corpus collision (53
rows, 32 wavs) needs `{index}_{stem}` IDs. (irv-ml1 `~/r16-vmoan-harness`)
- `[2026-06-11]` **A completion-poll `while pgrep -f <scriptname>` SELF-MATCHES its own remote shell
argv** — the poll's command line contains the script name, so its own `pgrep -f` always finds itself
→ the loop never exits, the poll never fires. Use a match pattern ABSENT from the poll command (pgrep
the python stage, or a sentinel file), not the driver's own name. (Caught only because the operator
asked "check status"; the job had already finished cleanly.)
- `[2026-06-08]` **Demucs `uv pip install demucs` pulls torch 2.12/torchaudio 2.11 → `ta.save()`
requires torchcodec → dies AFTER separating (0 stems written, rc=1).** Same class as the torch-2.12
torchcodec foot-gun. Fix = pin `torch==torchaudio==2.4.1` (pre-torchcodec save backend) +
`UV_LINK_MODE=copy` for the EPERM-hardlink quirk. Lesson restated: validate the SAVE path, not just
import + GPU inference, on a bleeding-edge torch.
- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — it fills the KV cache to the
`--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs
AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36
with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR `--enforce-eager`
(no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores
need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`)
- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres
rows (`dashboards`/`dashboard_widgets`); build by cloning a default-dashboard row + swapping the
measure. tok/s is NOT a per-generation field (null on the observation) — it's the
`outputTokensPerSecond` MEASURE, computed at metrics-API/dashboard query time; no native per-call
tok/s display exists (streaming doesn't change that). langfuse-web needs `HOSTNAME=0.0.0.0` (Next.js
standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host
3000 is gitea's → langfuse on 3001.
- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh +
corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read
world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are
vh-owned (editable), the `.env` is deploy-owned 600 (vh can't edit it, no sudo).
- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL` on
the containers; it defaults to claude-haiku in code, opt-in via config not `.env`. Don't trust an
".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection
corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.)
- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF —
the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the
model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline
-33pp type-F1 while valid_format held 1.0. An Ollama-matching `--chat-template` "fixed" it but was
the WRONG fix (the bundled scaffold is the divergent one). PRINCIPLE: serve each model's canonical
`tokenizer.apply_chat_template`, not the bundled template — bundled ones corrupt baselines. Verify
the applied prompt via vLLM `/tokenize`→`/detokenize`. (`90e08f0`/`27eb537`)
- `[2026-06-04]` **GPU pin by INDEX is ambiguous on irv-ml1** — native CUDA orders fastest-first
(A6000=0) but nvidia-smi/docker use PCI order (A6000=1), so an index pin can land on the wrong
card. Pin by **UUID** (`CUDA_VISIBLE_DEVICES=GPU-…`); verify via nvidia-smi compute-apps. Check
loaded-model VRAM with `ollama ps` (Ministral-3 @ its 256K default ctx = ~30 GB; cap num_ctx).
- `[2026-06-04]` **Persistent services on irv-ml1 need SYSTEM systemd units** — the box reaps
user-session processes on ssh disconnect, and `--user` systemd isn't reachable over non-login
ssh, so nohup/setsid/`screen -dmS`/`systemd-run --user` all die (even with `enable-linger`). Use
`/etc/systemd/system/`.
- `[2026-06-04]` **pyworld needs `setuptools<81`** (imports the removed `pkg_resources`); and
**R/soundgen `-lgfortran` fails** on irv-ml1 because the default `gcc` is gcc-11 but only
gfortran-12 is present (libgfortran.so lives only in the gcc-12 dir) → install `libgfortran-11-dev`.
- `[2026-06-04]` **homepage "crash" ≠ always NFS** — a wedged container in unkillable D-state
("tried to kill container, but did not receive an exit event") can come from dead `siteMonitor`
widget targets (retired ESH firewall IPs) hanging the node event loop into `exit_mmap`, needing a
host reboot. Check homepage's siteMonitors against retired hosts. (`incident_esh_docker_nfs_boot_race`)
_48 older entries archived to archival-memory.md._