380 lines
27 KiB
Markdown
380 lines
27 KiB
Markdown
# Persistent memory — eshpfi-management
|
||
|
||
_Last updated: 2026-06-13_
|
||
|
||
## Repo purpose
|
||
|
||
Reference workspace for PFI infrastructure: server inventory, canonical
|
||
Docker Compose stacks, ops playbooks, and conventions. Authoritative
|
||
copies of compose files live on the servers under
|
||
`/opt/docker/compose/<stack>/`; this repo mirrors them for version
|
||
control, editing, planning, and CI-driven deploys.
|
||
|
||
## Tools and conventions
|
||
|
||
Sister repos (separate gitea repos, deployed by playbooks here):
|
||
|
||
| Repo | Role | CI status |
|
||
|---|---|---|
|
||
| `vh/task-board` | MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
|
||
| `vh/vor` | Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
|
||
| `vh/nevermore` | Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
|
||
| `vh/asset-engine` | Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
|
||
| `vh/althing` | Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) | push-to-main → CI deploys (2026-05-14) |
|
||
| `vh/mead-hall` | Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
|
||
| `vh/skaldsong` | Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
|
||
| `vh/worldtree` | Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration | push-to-main → CI deploys |
|
||
| `vh/yt-voice-clipper` | YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → **gitea-webhook auto-deploy** to irv-ml1 (2026-06-03) — see `docs/runbooks/ytvc-autodeploy.md` |
|
||
|
||
(`vh/volva` + Heid were re-architected from systemd daemons to Claude Code
|
||
session orchestrators 2026-06-08; their nh3-dev `.service` units were removed —
|
||
no longer deployed sidecars here. See Recent decisions.)
|
||
|
||
- **Two-layer backups** — Backrest orchestrates restic for file+DB (5
|
||
fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for
|
||
VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore +
|
||
cross-site restic targets — see `docs/runbooks/disaster-recovery.md`
|
||
for the blast-radius matrix.
|
||
|
||
- **`pull-hf-repo.yaml`** is the canonical "get a HuggingFace
|
||
model/dataset onto ana-ml2's shared cache at
|
||
`/tank/aimodels/huggingface/`" playbook. Supports `--var repo_type=model|dataset|space`. Replaces ad-hoc `huggingface_hub.snapshot_download` patterns.
|
||
|
||
- **Worldtree admin auth — per-instance.** Each Worldtree deployment
|
||
(demo :8080, personal :8081, pinned :8082) has its own Heimdall
|
||
registry and its own bootstrap admin key. Infra-ops's stored
|
||
long-lived admin key (`key_id 61419c92`) at
|
||
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`
|
||
auths against **demo only**. For personal-instance admin ops, fetch
|
||
the bootstrap admin per-op via
|
||
`docker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEY`
|
||
on corviduo-dev. Used for `POST /admin/keys`, admin diagnostics
|
||
(`/admin/sessions/<id>/{bifrost,tools}`, etc.).
|
||
|
||
- **Per-project user keys against personal Worldtree** (issued
|
||
2026-05-19): `skaldsong:79744637` (nh3-dev iteration),
|
||
`skaldsong:7c1dbbbe` (ana-docker prod), `althing:50d85460`,
|
||
`mead-hall:a360822d`. Same `user_id=skaldsong` across both
|
||
skaldsong keys → shared Heimdall agent slot; different `key_id`
|
||
→ independently rotatable. Pattern: mint via `/admin/keys`, drop
|
||
value to `/tmp/wt-personal-<name>.key` mode 600, dev collects +
|
||
shreds (DO NOT cat to chat transcript).
|
||
|
||
- **Skaldsong CD pattern (registry-pull).** Differs from althing /
|
||
asset-engine which build-on-host. vh/skaldsong's CI builds and
|
||
pushes `gitea.phasefinal.com/vh/skaldsong:<sha>` + `:latest`;
|
||
`playbooks/deploy-skaldsong.yaml` on ana-docker pulls + recreates.
|
||
SHA-pin only (no `:latest` health-gated advance yet). Prereq: host
|
||
needs `docker login gitea.phasefinal.com` once (read:package PAT) —
|
||
not currently in the workflow.
|
||
|
||
- **docker-as-root pattern** (for ops that have no admin API, e.g.
|
||
`SqliteUserStore.set_bifrost_credentials`): on hosts where the SSH
|
||
user is in the `docker` group but lacks passwordless sudo, run
|
||
`docker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..."` to edit deploy-owned files
|
||
without sudo. Documented with security warning in
|
||
`servers/corviduo-dev/README.md`. docker-group membership is
|
||
effectively root via bind-mount; treat as a sudo-equivalent grant.
|
||
**Foot-gun: when running `docker compose` inside this sandbox,
|
||
any relative path in compose.yaml (e.g. `${WORLDTREE_CONFIG_DIR:-./config}`)
|
||
resolves against the sandbox CWD, but Docker daemon interprets the
|
||
resulting path against the HOST filesystem. Always pass `-e VAR=/abs/path`
|
||
to the docker run invocation for any relative-default config dir.**
|
||
|
||
- **`scripts/elway` sudo handling** — elway prompts for the sudo password
|
||
ONCE via `getpass` before the first `sudo: true` step. That prompt is
|
||
interactive → elway can't run unattended from a non-TTY tool if any step
|
||
needs sudo. For sudo-free playbooks (no `sudo: true` steps) it runs fully
|
||
non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo,
|
||
use the docker-daemon-root trick: `docker run --rm -v /worktank:/mnt alpine
|
||
sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'`.
|
||
|
||
## Current state / in-flight
|
||
|
||
_As of 2026-06-13:_
|
||
|
||
- **ana-ml2 went Ada → Blackwell** (dual NVIDIA RTX PRO 6000 Blackwell Max-Q, 96 GB each, cc 12.0 /
|
||
sm_120 — was dual RTX 6000 Ada 48 GB / cc 8.9, confirmed live via `nvidia-smi`). GPU layout now:
|
||
**GPU 0** held free for large-model hot-loads (llama-swap pinned there, `edf0f91`); **GPU 1** is the
|
||
steady-tenant card — granite (131k ctx) + qwen3.5-vl (65k) + the embed/rerank/reward trio, ~3.5 GB
|
||
free after the rebalance. CLAUDE.md's server table + the granite/qwen "Ada (cc 8.9)" compose comments
|
||
are STALE — doc-fix offered, **pending operator go-ahead** (/snapshot doesn't auto-edit CLAUDE.md).
|
||
|
||
- **R16 "vmoan" TTS LoRA POC — awaiting the operator's ear** (Brokkr's designated gate). Two auditions
|
||
on NFS: `/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v1.wav` (42 s, run-on) vs `_v2.wav` (10 s,
|
||
tight). v2 fixed the run-on; Brokkr holds the pilot verdict on the operator's listen. Grab:
|
||
`scp irv-ml1:/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v{1,2}.wav .`. LOCAL-ONLY corpus
|
||
(soundgasm-derived) — distribution barred; never echo transcripts to the bus.
|
||
|
||
- **R17 v2 corpus characterization still on the local irv-ml1 branch, push HELD** (`r17-v2-character
|
||
ization`, 18 MULTI / 101 ELIGIBLE + 19936 candidates + `speech.json` ASR). No-push rule AND ASR of
|
||
the intimate-audio batch (content-exposure call). Awaiting brokkr-collect or explicit push approval.
|
||
|
||
- **Mac Pro migration is still the big open project** — `migration-plan.md` (repo root): WORKSTATION-
|
||
ONLY move of nh3-dev's Hat-1 dev env (~40 repos, all of `~/.claude`, dotfiles, toolchain) to an M2
|
||
Ultra Mac Pro Rack on the 10.100 subnet. Hat-2 fleet sidecars (egress SOCKS5, mead-hall, ttyd seats,
|
||
`althing-forseti`) STAY on the Linux VM. **Phase 0** (push pushables + confirm the ~6 local-only repos)
|
||
is the only time-sensitive step. Cutover = `rsync` working trees, NOT re-clone (re-clone loses unsaved
|
||
work in ~15 repos). macOS `/home/lkraven`→`/Users/lkraven` repath. Everything else waits on hardware.
|
||
|
||
- **sglang-vs-vLLM bench stack staged but parked** (`5f049cb`) — the originating question (SGLang
|
||
RadixAttention caching) resolved without it: vLLM v1 already defaults prefix-caching ON, now pinned
|
||
explicit on granite+qwen. The NVFP4 chase is abandoned. Bench stack stays staged if ever revisited;
|
||
the MAX_JOBS-on-shared-prod foot-gun (see Tried) caps any future from-source build on ana-ml2.
|
||
|
||
- **pi + GLM 5.1 harness live on nh3-dev** — `glm` launcher runs pi against `glm-5.1` (thinking-off)
|
||
via the LiteLLM gateway; `glm-5.1-reasoning` for opt-in thinking.
|
||
|
||
- **Worldtree config-propagation lane is mature + humming** — pre-merge delta-ping → sync-on-merge to
|
||
the demo+personal bind-mounts; pinned stays pre-cutover (full migration at its next re-image).
|
||
|
||
- **Disclosed-keys hygiene queue** (rotate at convenience): HF token, the wt-personal keys, Gitea
|
||
runner reg token, MINIFLUX_PASSWORD, ZAI_API_KEY. (The shared all-agents LiteLLM key is intentional,
|
||
not hygiene-debt — see decisions; rotatable via infra-ops only if it leaks.)
|
||
|
||
- **Still open from prior:** clean legacy `news-digest` on ana-docker; the `docker push 60s ceiling`
|
||
mystery uninstrumented. (llama-swap GPU-0 pin — DONE this session, `edf0f91`.)
|
||
|
||
## Recent decisions
|
||
|
||
- `[2026-06-13]` **ana-ml2 upgraded Ada → dual RTX PRO 6000 Blackwell Max-Q** (96 GB each, cc 12.0 /
|
||
sm_120; was dual RTX 6000 Ada 48 GB / cc 8.9 — confirmed live via `nvidia-smi`). Unlocks NVFP4 (FP4
|
||
tensor cores) and doubles VRAM headroom. CLAUDE.md's server table + the granite/qwen compose
|
||
"Ada (cc 8.9)" comments are now stale; doc-fix **offered, pending operator go-ahead** (snapshot doesn't
|
||
auto-edit CLAUDE.md). Tracking surface: this snapshot + commits `19a07b9`/`1e2a3a1` ("Blackwell 96GB").
|
||
|
||
- `[2026-06-13]` **NVFP4-W4A4 is infeasible for Granite — FP8 stays the Granite-on-Blackwell format.**
|
||
W4A4 (4-bit weights + activations) collapses at 30k context, proven **producer-independent** (modelopt
|
||
AND llm-compressor both clean-NONE from the same BF16 base + wikitext-2k calib). No 4-bit wins both
|
||
axes: W4A4 = quality collapse; W4A16-NVFP4/AWQ = weight-only dequant → bf16 (no FP4-core speedup), so
|
||
the tempting ~40% uplift doesn't materialize without quality loss. **30B retired** (not enough quality
|
||
for the VRAM/perf hit for our use case). (auto-memory `reference_nvfp4_w4a4_granite_infeasible`)
|
||
|
||
- `[2026-06-13]` **Qwen3.5-9B VL (FP8) deployed on ana-ml2 GPU 1** — `qwen35-vl` stack, :8007, gateway
|
||
alias `qwen3.5-9b-fp8`. Clean FP8 (no NVFP4 for vision). **Pinned nightly digest, not `:latest`**: the
|
||
stable release quantizes the Qwen3.5-VL *vision tower* under `--quantization fp8` → garbage vision (LM
|
||
fine, "sees" noise); the nightly correctly excludes the vision tower. Re-pin to `:latest` + drop the
|
||
pin once that exclusion lands stable. (`2e3dcc2`)
|
||
|
||
- `[2026-06-13]` **comfyui 325 G model tree migrated worktank → `/storetank/arbo`** (worktank 97% → 26%).
|
||
`arbo` is the consuming app (ComfyUI-backed prompt/image pipeline); models live under its name on the
|
||
bigger pool. Overlay bind-mount via `COMFYUI_MODELS_DIR` in the comfyui compose; inventory at
|
||
`docs/arbo-comfyui-model-catalog.md` (retain-large-portion-for-future-use decision). (`38186be`)
|
||
|
||
- `[2026-06-13]` **GPU layout settled on the Blackwell box.** GPU 0 held free for large-model hot-loads
|
||
(llama-swap pinned there, `edf0f91`); GPU 1 is the steady-tenant card — granite 131k ctx (was 51k),
|
||
qwen 65k, embed/rerank/reward trio, ~3.5 GB free after rebalance (`1e2a3a1`, `19a07b9`; trio re-floored
|
||
for 96 GB, 20×-parallel-stable). embed/rerank left at floor — long docs are chunked BEFORE embedding,
|
||
so a longer embed ctx buys nothing. PagedAttention note: max-model-len is a ceiling, not a reservation,
|
||
so small requests aren't blocked by the big ceiling; concurrency = KV-pool-tokens / actual-request-size.
|
||
|
||
- `[2026-06-13]` **Prefix caching pinned explicit on granite + qwen** — benched ~6.5× faster TTFT (45 ms
|
||
vs 292 ms) on a shared ~4.5k-token summarizer template; soft/evictable KV, neutral when prefixes don't
|
||
repeat. vLLM v1 `:latest` defaults it ON (granite) but the qwen nightly defaults it OFF — pin both so a
|
||
version flip can't silently disable it. (`a9a2be7`)
|
||
|
||
- `[2026-06-13]` **granite-4.1-8b listed as the always-available summarizer/classifier endpoint + a
|
||
shared all-agents key minted** (operator-directed). Added to the global `~/.claude/CLAUDE.md` Global-
|
||
tools section; key alias `all-agents-local`, scoped to the FREE local models only (granite + qwen-
|
||
vision + embed/rerank, NOT the paid GLM), internal-gateway-only, rotatable. The standing "reach for
|
||
this before spending premium tokens on low-caliber high-volume work" lever. (auto-memory
|
||
`reference_litellm_gateway`)
|
||
|
||
- `[2026-06-13]` **arbo engine + frontend stack stood up** (ADR-0001) — irv-ml1 co-located inference
|
||
engine (`ee57e69`), python-based healthcheck (the slim image ships no curl/wget, `bdb3312`), frontend
|
||
ro-mounted from the v0.11.2 checkout (`922e8ad`, ADR-0001 D2). arbo = the ComfyUI-consuming app whose
|
||
models now live at `/storetank/arbo`.
|
||
|
||
- `[2026-06-11]` **GLM thinking inverted at the LiteLLM gateway** (operator call): `glm-5.1` defaults
|
||
thinking-OFF; new `glm-5.1-reasoning` alias = same z.ai upstream with thinking ON (opt-in).
|
||
Mechanism: `litellm_params.extra_body:{thinking:{type:disabled}}` — LiteLLM `drop_params` STRIPS a
|
||
top-level `thinking`/`reasoning_effort`, but forwards `extra_body` verbatim to z.ai (the only channel
|
||
that works; verified reasoning_tokens 0 vs >0). Shared-gateway change — affects ALL glm-5.1 callers
|
||
(brokkr's all-models key included). (`95b2701`, auto-memory `reference_litellm_gateway`)
|
||
|
||
- `[2026-06-11]` **pi coding agent installed on nh3-dev as a GLM 5.1 harness** — `@earendil-works/
|
||
pi-coding-agent` via **bun** (user-level; npm's global prefix is `/usr` → needs sudo, bun avoids it).
|
||
Config `~/.pi/agent/models.json` (litellm provider → gateway), launcher `~/.local/bin/glm` sources
|
||
the gateway key + selects the model. pi is OpenAI-compatible; proxy-safe compat flags for the GLM path.
|
||
|
||
- `[2026-06-11]` **z.ai web-tools (regin) = z.ai hosted MCP path, NOT the `/paas/v4` Tool API.** Direct
|
||
`api.z.ai/api/paas/v4/web_search` → 429/1113 "insufficient balance" (coding-plan keys bill tools on a
|
||
separate quota path); `open.bigmodel.cn` is the China platform (404, different account). WORKS: MCP
|
||
streamable-HTTP at `https://api.z.ai/api/mcp/{web_search_prime,web_reader}/mcp`, `Authorization: Bearer
|
||
$ZAI_API_KEY` (the **MCP** key — distinct from `Z_AI_API_KEY` the LLM key). Reference impl = Worldtree's
|
||
Leif agent (`tools/web/zai_client.py` + `core/clients/mcp.py`).
|
||
|
||
- `[2026-06-10]` **Mac Pro migration framed: workstation-only** (M2 Ultra ARM, racked NH3 on-subnet);
|
||
sidecars stay Linux. `migration-plan.md`. (See in-flight.)
|
||
|
||
- `[2026-06-10]` **Worldtree deployed-config propagation is infra-ops's OWNED lane** (operator ruling).
|
||
worldtree-dev pings the config delta on every config-touching commit (pre-merge); infra-ops syncs
|
||
`config/*.yaml` from MERGED canonical to the `/opt/worldtree*/config` bind-mounts on demo+personal
|
||
(Heimdall hot-reloads policies; defaults are code-defaulted). The **v0.33.8 9-HOUR demo outage** — a
|
||
`model_roles.yaml` hard-startup-dep that shipped in canonical 3 releases earlier but never reached the
|
||
VM — is the failure mode this lane prevents. providers.yaml stays hand-tuned (artemis graft on
|
||
personal). corviduo-dev emergency-ops = `ssh vh@10.250.50.152` (alias doesn't resolve, lkraven denied,
|
||
infra-ops key excluded), docker no-sudo. (auto-memory `reference_worldtree_deploys_cicd`,
|
||
`reference_corviduo_dev_emergency_ops`)
|
||
|
||
- `[2026-06-09]` **LiteLLM scoped virtual keys issued to consumers** (operator-authorized): `brokkr-
|
||
smithy` (all-proxy-models), `arbo-prompt-enhance` (comfy-dev, granite-only). Pattern: mint via
|
||
`/key/generate` (master `sk-corvid`), scope-restricted + rotatable, value → 600 file never the bus.
|
||
(auto-memory `reference_litellm_gateway`)
|
||
|
||
- `[2026-06-08]` **volva.service + heid.service removed from nh3-dev** — vestigial systemd daemons;
|
||
Heid/Volva re-architected from Python pollers to Claude Code session orchestrators (heid commit
|
||
`12aa5a9`); binaries+dirs gone, volva.service was crash-looping 203/EXEC. Cleanup at heid's request
|
||
(the bus-content-driven sudo got the harness guardrail; operator green-lit). (`6e2f80e`)
|
||
|
||
- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer** (supersedes
|
||
the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not
|
||
Q4/AWQ** — primary workload (agent memory + summarization) is high-concurrency, where FP8 scales
|
||
~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a config bug).
|
||
vLLM `vllm-granite` :8004 GPU 1, official IBM compressed-tensors FP8, CUDA graphs. (Then on Ada cc 8.9;
|
||
the box has since gone Blackwell — see 2026-06-13.) (`34a43a0`, auto-memory `reference_ana_ml2_vllm_granite`)
|
||
|
||
- `[2026-06-05]` **Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI**; LiteLLM
|
||
`success_callback:[langfuse]` live (project `gateway`). Pretty prompt/completion/reasoning traces +
|
||
an `outputTokensPerSecond` tok/s dashboard. NOT a prerequisite — spend_logs already capture
|
||
tokens+latency. (`9171e6a`, auto-memory `reference_litellm_gateway`)
|
||
|
||
- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any
|
||
found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memory
|
||
`feedback_avoid_ollama`)
|
||
|
||
- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent).
|
||
FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI;
|
||
eshpfi keeps the `comfyui` stack compose, comfy-dev owns the model/workflow knowledge. (auto-memory
|
||
`reference_irv_ml1_ampere_quant`)
|
||
|
||
- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is
|
||
the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to
|
||
claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands
|
||
the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts.
|
||
**CORRECTION to the 2026-06-04 "deploys ALL CICD" line:** the bind-mount CONFIGS (providers.yaml,
|
||
vh-owned on corviduo `/opt/worldtree*/config`) ARE infra-ops's to apply directly — only the
|
||
app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`)
|
||
|
||
- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent;
|
||
granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from
|
||
Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`)
|
||
|
||
- `[2026-06-04]` **phi4 ships the CANONICAL/official Phi-4 chat template, NOT Ollama's.**
|
||
Ollama's bundled template omits the system `<|end|>` — that flattered brokkr's R15 eval but is
|
||
the DIVERGENT scaffold (Dvalin: the system `<|end|>` is Microsoft's intended format). Applied an
|
||
Ollama-matching override then reverted — ship correct, not the benchmark quirk. (`90e08f0`→`27eb537`;
|
||
"headgun" lesson in Tried.)
|
||
|
||
- `[2026-06-04]` **infra-ops NOPASSWD-sudo identity commissioned, scoped to PFI boxes** (+esh-docker-vm
|
||
by operator override) — so infra-ops completes DevOps end-to-end vs handing the operator sudo steps.
|
||
Dedicated key, sudo log_output, key-gated. (`8c32a05`)
|
||
|
||
- `[2026-06-04]` **Worldtree demo/pinned/personal deploys are ALL CI/CD, not infra-ops** — a "deploy
|
||
vX.Y.Z" request to infra-ops is MISROUTED → point them back to their pipeline. The granite→phi4
|
||
repoint: worldtree-dev self-served via their CI/CD (v0.30.10). (`d8d776c`)
|
||
|
||
- `[2026-06-04]` **ollama upgraded 0.9.0→0.30.4 on irv-ml1** (Ministral-3 is a Dec-2025 model the
|
||
old engine refused); A6000 pinned by **UUID** not index (native fastest-first ≠ nvidia-smi PCI).
|
||
|
||
- `[2026-06-04]` **`brokkr` user (no-sudo) on irv-ml1; R14/R15/R16 substrate moved to /home/brokkr.**
|
||
Persistent box services there need SYSTEM systemd units (see Tried).
|
||
|
||
_44 older entries archived to archival-memory.md._
|
||
|
||
## Tried and abandoned
|
||
|
||
- `[2026-06-13]` **Heavy from-source compile (`MAX_JOBS=128` sglang fork build) on the shared PROD GPU
|
||
box PINS it** — load hit 187, prod vLLM services restarted, killed an in-flight quant. ana-ml2 hosts
|
||
live inference; never run a big build there at full parallelism. Cap `MAX_JOBS≤32`, build off-box, or
|
||
cgroup-constrain. (Operator ran the kill; infra-ops NOPASSWD-sudo confirmed working on ana-ml2 — retry
|
||
infra-ops on an ssh-255 before concluding "no access.")
|
||
|
||
- `[2026-06-13]` **`--quantization fp8` on a VL model can quantize the VISION TOWER → garbage vision**
|
||
(Qwen3.5-VL on the stable vLLM: gray-grid output; the LM answers text fine, so it "looks" healthy
|
||
until you actually feed it an image). The nightly excludes the vision tower. Lesson: validate the
|
||
VISION path on a quantized VLM, not just text — and pin the engine digest that has the exclusion.
|
||
|
||
- `[2026-06-13]` **vLLM's `--gpu-memory-utilization` is checked against FREE VRAM at startup
|
||
(`free >= util*total`), not total** — on a shared card, growing one service before trimming a
|
||
co-tenant OOMs ("free 48.56 < desired 80.72" at util 0.85 on a half-occupied 96 GB card). Start-order
|
||
matters: trim the shrinking service FIRST, then grow the other. Size to the FREE budget, not the total.
|
||
|
||
- `[2026-06-13]` **The `vllm/vllm-openai` entrypoint is already `["vllm","serve"]`** — the compose
|
||
`command:` supplies the model as the first POSITIONAL arg + flags; a second `serve` (or `--model X`)
|
||
yields "unrecognized arguments". Same-class gotchas this session: `tee` masks the real exit code (use
|
||
a `>` redirect to keep rc); HF `datasets` rejects bare `wikitext` (needs `Salesforce/wikitext`).
|
||
|
||
- `[2026-06-13]` **Chatterbox-Turbo LoRA finetune: the repo's `setup.py` loads the WRONG tokenizer** —
|
||
`merge_and_save_turbo_tokenizer()` pulls gpt2-medium + a grapheme merge file (len mismatch) instead of
|
||
the chatterbox-turbo GPT2 tokenizer (vocab.json+merges.txt, len 50276). Fix = override with the correct
|
||
tokenizer + delete the grapheme `tokenizer.json`; `[vmoan]` token → new_vocab_size 50277 (1-row resize),
|
||
lora_r 64 / alpha 128, modules_to_save=[text_emb,text_head]. Also: a unique-stem corpus collision (53
|
||
rows, 32 wavs) needs `{index}_{stem}` IDs. (irv-ml1 `~/r16-vmoan-harness`)
|
||
|
||
- `[2026-06-11]` **A completion-poll `while pgrep -f <scriptname>` SELF-MATCHES its own remote shell
|
||
argv** — the poll's command line contains the script name, so its own `pgrep -f` always finds itself
|
||
→ the loop never exits, the poll never fires. Use a match pattern ABSENT from the poll command (pgrep
|
||
the python stage, or a sentinel file), not the driver's own name. (Caught only because the operator
|
||
asked "check status"; the job had already finished cleanly.)
|
||
|
||
- `[2026-06-08]` **Demucs `uv pip install demucs` pulls torch 2.12/torchaudio 2.11 → `ta.save()`
|
||
requires torchcodec → dies AFTER separating (0 stems written, rc=1).** Same class as the torch-2.12
|
||
torchcodec foot-gun. Fix = pin `torch==torchaudio==2.4.1` (pre-torchcodec save backend) +
|
||
`UV_LINK_MODE=copy` for the EPERM-hardlink quirk. Lesson restated: validate the SAVE path, not just
|
||
import + GPU inference, on a bleeding-edge torch.
|
||
|
||
- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — it fills the KV cache to the
|
||
`--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs
|
||
AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36
|
||
with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR `--enforce-eager`
|
||
(no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores
|
||
need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`)
|
||
|
||
- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres
|
||
rows (`dashboards`/`dashboard_widgets`); build by cloning a default-dashboard row + swapping the
|
||
measure. tok/s is NOT a per-generation field (null on the observation) — it's the
|
||
`outputTokensPerSecond` MEASURE, computed at metrics-API/dashboard query time; no native per-call
|
||
tok/s display exists (streaming doesn't change that). langfuse-web needs `HOSTNAME=0.0.0.0` (Next.js
|
||
standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host
|
||
3000 is gitea's → langfuse on 3001.
|
||
|
||
- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh +
|
||
corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read
|
||
world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are
|
||
vh-owned (editable), the `.env` is deploy-owned 600 (vh can't edit it, no sudo).
|
||
|
||
- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL` on
|
||
the containers; it defaults to claude-haiku in code, opt-in via config not `.env`. Don't trust an
|
||
".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection
|
||
corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.)
|
||
|
||
- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF —
|
||
the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the
|
||
model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline
|
||
-33pp type-F1 while valid_format held 1.0. An Ollama-matching `--chat-template` "fixed" it but was
|
||
the WRONG fix (the bundled scaffold is the divergent one). PRINCIPLE: serve each model's canonical
|
||
`tokenizer.apply_chat_template`, not the bundled template — bundled ones corrupt baselines. Verify
|
||
the applied prompt via vLLM `/tokenize`→`/detokenize`. (`90e08f0`/`27eb537`)
|
||
|
||
- `[2026-06-04]` **GPU pin by INDEX is ambiguous on irv-ml1** — native CUDA orders fastest-first
|
||
(A6000=0) but nvidia-smi/docker use PCI order (A6000=1), so an index pin can land on the wrong
|
||
card. Pin by **UUID** (`CUDA_VISIBLE_DEVICES=GPU-…`); verify via nvidia-smi compute-apps. Check
|
||
loaded-model VRAM with `ollama ps` (Ministral-3 @ its 256K default ctx = ~30 GB; cap num_ctx).
|
||
|
||
- `[2026-06-04]` **Persistent services on irv-ml1 need SYSTEM systemd units** — the box reaps
|
||
user-session processes on ssh disconnect, and `--user` systemd isn't reachable over non-login
|
||
ssh, so nohup/setsid/`screen -dmS`/`systemd-run --user` all die (even with `enable-linger`). Use
|
||
`/etc/systemd/system/`.
|
||
|
||
- `[2026-06-04]` **pyworld needs `setuptools<81`** (imports the removed `pkg_resources`); and
|
||
**R/soundgen `-lgfortran` fails** on irv-ml1 because the default `gcc` is gcc-11 but only
|
||
gfortran-12 is present (libgfortran.so lives only in the gcc-12 dir) → install `libgfortran-11-dev`.
|
||
|
||
- `[2026-06-04]` **homepage "crash" ≠ always NFS** — a wedged container in unkillable D-state
|
||
("tried to kill container, but did not receive an exit event") can come from dead `siteMonitor`
|
||
widget targets (retired ESH firewall IPs) hanging the node event loop into `exit_mmap`, needing a
|
||
host reboot. Check homepage's siteMonitors against retired hosts. (`incident_esh_docker_nfs_boot_race`)
|
||
|
||
_48 older entries archived to archival-memory.md._
|