diff --git a/persistent-memory.md b/persistent-memory.md index fd8ae7a..6224f71 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-06-04_ +_Last updated: 2026-06-05_ ## Repo purpose @@ -88,50 +88,60 @@ Sister repos (separate gitea repos, deployed by playbooks here): ## Current state / in-flight -_As of 2026-06-04:_ +_As of 2026-06-05:_ -- **INFRA SESSION 2026-06-04 — phi4 on vLLM, infra-ops sudo identity, R15/R16, brokkr svc.** - Major threads (detail in cited auto-memories + commits): - - **phi4-mini FP8 LIVE on ana-ml2 vLLM** (`vllm-phi4`, :8004, GPU 1, 50K ctx, FP8 + FP8-KV) - as the **nevermore** summarizer/dreaming agent — superseded **granite-4-small** (removed - from llama-swap config; GGUFs kept on disk). Template = **CANONICAL/official Phi-4** (final, - after an Ollama-matching override applied `90e08f0` then reverted `27eb537`). nevermore - repointed (LLAMA_SWAP_URL→:8004, MODEL→phi4-mini). auto-memory `reference_ana_ml2_vllm_phi4`. - OPEN: brokkr re-baselining R15 P02 under canonical (eval re-regresses ~0.78→~0.45-0.57; not - live — worldtree #252 is baseline-first); **GPU 1 tight (~10 GB free — pin llama-swap to - GPU 0 as follow-up)**; Worldtree Vili #253 (hardcoded :9292 granite fallback, their fix). - - **infra-ops NOPASSWD-sudo identity commissioned** across PFI boxes — `ssh infra-ops@`, - key `~/.ssh/infra-ops_ed25519`. Bootstrap `playbooks/bootstrap-infra-ops-user.yaml` + - `scripts/bootstrap-infra-ops-fleet.sh` (tiers 1+2 live: irv-ml1/ana-ml2/ana-docker/nh3-docker/ - ana-nas + 4 app VMs; **esh-docker-vm added by operator override**). Excludes sf-*/corviduo/ - Synology. auto-memory `reference_infra_ops_sudo_identity` (`8c32a05`). - - **R15/R16 brokkr-smithy harness stood up on irv-ml1** — ollama upgraded **0.9.0→0.30.4** - (Ministral-3 needs it), A6000 **UUID-pinned**; R15 ollama arms (granite4.1:3b/8b, qwen3:4b, - phi4-mini:3.8b, hf SmolLM3-GGUF, ministral-3:3b-instruct, nuextract:3.8b) + R16 R/soundgen + - pyworld venv. `playbooks/irv-ml1-r15-r16-{nosudo,sudo}.yaml`. auto-memory `reference_irv_ml1_gpu_r14`. - - **`brokkr` user created on irv-ml1** (no-sudo) + 24 GB R14/R15/R16 substrate migrated out of - lkraven's home → `/home/brokkr/`; gitea pull = read-only deploy key. **brokkr-audition.service** - (SYSTEM systemd unit, :8137) serves brokkr's R16 NVV audition UI (the morph set being auditioned). - - **homepage incident (esh-docker-vm)** — wedged on dead siteMonitor IP (retired ESH firewall - 10.0.250.1) into unkillable D-state; host reboot cleared it; ESH-Firewall widget removed from - `services.yaml`. (`incident_esh_docker_nfs_boot_race` updated.) - - **Observability roadmap** `docs/roadmap.md` — Langfuse (full req/resp tracing) + - Prometheus/Grafana off vLLM `/metrics`. Deferred; Langfuse first. - -- **Disclosed-keys hygiene queue** — rotate at convenience: HF token `hf_HBl…` - (lkraven's HF account) leaked into BuildKit logs during the CSM build attempt - (logs shredded, never committed — low urgency); - `/tmp/wt-personal-skaldsong-prod.key` on nh3-dev; mead-hall's prior Worldtree - bearer (superseded by `a360822d`); Worldtree `Z_AI_API_KEY`/`ZAI_API_KEY`; - chamber `forseti`/`agent_runner` api_keys (superseded by `50d85460`); Gitea - runner registration token (`a1135753…`). -- **Still open from prior sessions:** rotate `MINIFLUX_PASSWORD` (leaked - twice); clean up legacy `news-digest` detritus on ana-docker; watch nh3-nas - `/volume1` (was 65%; recheck before ~80%); the `docker push 60s client-side - ceiling` mystery remains uninstrumented. +- **Granite-FP8 + observability session — all LIVE & committed (`34a43a0`, `9171e6a`).** + - **Granite 4.1 8B FP8 is the production summarizer** (`vllm-granite` :8004, ana-ml2 GPU 1, + 50K ctx, CUDA graphs) — replaced phi4-mini, validated by brokkr (valid_format 1.0, FP8 stays). + - **LiteLLM gateway** (:4000) routes `granite-4.1-8b`→vLLM (explicit entry shadows the `*` + wildcard) + **Langfuse v3 wired** (ana-docker:3001, "LLM Throughput (tok/s)" dashboard built). + - **GPU-1 retuned** (trio over-provisioned KV trimmed) → granite runs with CUDA graphs + ~10 GB + free as a future Granite-text-LoRA hedge. Streaming through the gateway confirmed (TTFT 0.24s). + - **ana-docker pruned** 77 GB (unused images + build cache; disk 83%→49%) to fit ClickHouse. +- **Worldtree summarizer repoint — NO instance change now; DEFERRED to Worldtree #254** (see Recent + decisions). worldtree-dev will ping with the providers.yaml + consumer config when #254 un-holds; + infra-ops applies to the personal/demo/pinned bind mounts (vh@10.250.50.152, `/opt/worldtree*/config`). +- **Commits unpushed** (`34a43a0`, `9171e6a`, nevermore `d3e19b8` in its repo) — operator's call to push. +- **Operator flagged "new work to do"** for the next session — this snapshot is the handoff. +- **Disclosed-keys hygiene queue** (rotate at convenience): HF token `hf_HBl…` (lkraven's), `/tmp/ + wt-personal-skaldsong-prod.key`, Worldtree `Z_AI_API_KEY`, Gitea runner reg token, `MINIFLUX_PASSWORD` + (leaked twice). (sk-corvid + the langfuse/vastblueai-gateway keys are dev-enclosed — leakage deprioritized.) +- **Still open from prior:** clean legacy `news-digest` on ana-docker; watch nh3-nas `/volume1`; **pin + llama-swap to GPU 0** for clean GPU-1 separation; the `docker push 60s ceiling` mystery uninstrumented. ## Recent decisions +- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer** (supersedes + the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not + Q4/AWQ** — primary workload (agent memory + summarization) is high-concurrency, where FP8-on-Ada + scales ~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a + config bug — placement/kernel/contention all ruled out). vLLM `vllm-granite` :8004 GPU 1, official + IBM compressed-tensors FP8, CUDA graphs. **GPU-1 retune** (trio utils 0.2/0.2/0.3→0.07/0.07/0.18, + granite 0.36) freed ~10 GB → CUDA graphs + a Granite-text-LoRA hedge. nevermore repointed. (`34a43a0`, + auto-memory `reference_ana_ml2_vllm_granite`) + +- `[2026-06-05]` **Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI**; LiteLLM + `success_callback:[langfuse]` live (project `gateway`). Pretty prompt/completion/reasoning traces + + an `outputTokensPerSecond` tok/s dashboard. NOT a prerequisite — spend_logs already capture + tokens+latency. (`9171e6a`, auto-memory `reference_litellm_gateway`) + +- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any + found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memory + `feedback_avoid_ollama`) + +- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent). + FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI; + eshpfi keeps the `comfyui` stack compose, comfy-dev owns the model/workflow knowledge. (auto-memory + `reference_irv_ml1_ampere_quant`) + +- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is + the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to + claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands + the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts. + **CORRECTION to the 2026-06-04 "deploys ALL CICD" line:** the bind-mount CONFIGS (providers.yaml, + vh-owned on corviduo `/opt/worldtree*/config`) ARE infra-ops's to apply directly — only the + app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`) + - `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent; granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`) @@ -206,6 +216,31 @@ _37 older entries archived to archival-memory.md._ ## Tried and abandoned +- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — it fills the KV cache to the + `--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs + AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36 + with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR `--enforce-eager` + (no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores + need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`) + +- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres + rows (`dashboards`/`dashboard_widgets`); build by cloning a default-dashboard row + swapping the + measure. tok/s is NOT a per-generation field (null on the observation) — it's the + `outputTokensPerSecond` MEASURE, computed at metrics-API/dashboard query time; no native per-call + tok/s display exists (streaming doesn't change that). langfuse-web needs `HOSTNAME=0.0.0.0` (Next.js + standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host + 3000 is gitea's → langfuse on 3001. + +- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh + + corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read + world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are + vh-owned (editable), the `.env` is deploy-owned 600 (vh can't edit it, no sudo). + +- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL` on + the containers; it defaults to claude-haiku in code, opt-in via config not `.env`. Don't trust an + ".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection + corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.) + - `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF — the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline