memory: snapshot — 2026-06-05 granite-FP8 cutover + Langfuse observability + worldtree #254 deferral

This commit is contained in:
vh
2026-06-05 14:14:44 -07:00
parent 9171e6a20f
commit 5349da567c
+76 -41
View File
@@ -1,6 +1,6 @@
# Persistent memory — eshpfi-management
_Last updated: 2026-06-04_
_Last updated: 2026-06-05_
## Repo purpose
@@ -88,50 +88,60 @@ Sister repos (separate gitea repos, deployed by playbooks here):
## Current state / in-flight
_As of 2026-06-04:_
_As of 2026-06-05:_
- **INFRA SESSION 2026-06-04 — phi4 on vLLM, infra-ops sudo identity, R15/R16, brokkr svc.**
Major threads (detail in cited auto-memories + commits):
- **phi4-mini FP8 LIVE on ana-ml2 vLLM** (`vllm-phi4`, :8004, GPU 1, 50K ctx, FP8 + FP8-KV)
as the **nevermore** summarizer/dreaming agent — superseded **granite-4-small** (removed
from llama-swap config; GGUFs kept on disk). Template = **CANONICAL/official Phi-4** (final,
after an Ollama-matching override applied `90e08f0` then reverted `27eb537`). nevermore
repointed (LLAMA_SWAP_URL→:8004, MODEL→phi4-mini). auto-memory `reference_ana_ml2_vllm_phi4`.
OPEN: brokkr re-baselining R15 P02 under canonical (eval re-regresses ~0.78→~0.45-0.57; not
live — worldtree #252 is baseline-first); **GPU 1 tight (~10 GB free — pin llama-swap to
GPU 0 as follow-up)**; Worldtree Vili #253 (hardcoded :9292 granite fallback, their fix).
- **infra-ops NOPASSWD-sudo identity commissioned** across PFI boxes — `ssh infra-ops@<host>`,
key `~/.ssh/infra-ops_ed25519`. Bootstrap `playbooks/bootstrap-infra-ops-user.yaml` +
`scripts/bootstrap-infra-ops-fleet.sh` (tiers 1+2 live: irv-ml1/ana-ml2/ana-docker/nh3-docker/
ana-nas + 4 app VMs; **esh-docker-vm added by operator override**). Excludes sf-*/corviduo/
Synology. auto-memory `reference_infra_ops_sudo_identity` (`8c32a05`).
- **R15/R16 brokkr-smithy harness stood up on irv-ml1** — ollama upgraded **0.9.0→0.30.4**
(Ministral-3 needs it), A6000 **UUID-pinned**; R15 ollama arms (granite4.1:3b/8b, qwen3:4b,
phi4-mini:3.8b, hf SmolLM3-GGUF, ministral-3:3b-instruct, nuextract:3.8b) + R16 R/soundgen +
pyworld venv. `playbooks/irv-ml1-r15-r16-{nosudo,sudo}.yaml`. auto-memory `reference_irv_ml1_gpu_r14`.
- **`brokkr` user created on irv-ml1** (no-sudo) + 24 GB R14/R15/R16 substrate migrated out of
lkraven's home → `/home/brokkr/`; gitea pull = read-only deploy key. **brokkr-audition.service**
(SYSTEM systemd unit, :8137) serves brokkr's R16 NVV audition UI (the morph set being auditioned).
- **homepage incident (esh-docker-vm)** — wedged on dead siteMonitor IP (retired ESH firewall
10.0.250.1) into unkillable D-state; host reboot cleared it; ESH-Firewall widget removed from
`services.yaml`. (`incident_esh_docker_nfs_boot_race` updated.)
- **Observability roadmap** `docs/roadmap.md` — Langfuse (full req/resp tracing) +
Prometheus/Grafana off vLLM `/metrics`. Deferred; Langfuse first.
- **Disclosed-keys hygiene queue** — rotate at convenience: HF token `hf_HBl…`
(lkraven's HF account) leaked into BuildKit logs during the CSM build attempt
(logs shredded, never committed — low urgency);
`/tmp/wt-personal-skaldsong-prod.key` on nh3-dev; mead-hall's prior Worldtree
bearer (superseded by `a360822d`); Worldtree `Z_AI_API_KEY`/`ZAI_API_KEY`;
chamber `forseti`/`agent_runner` api_keys (superseded by `50d85460`); Gitea
runner registration token (`a1135753…`).
- **Still open from prior sessions:** rotate `MINIFLUX_PASSWORD` (leaked
twice); clean up legacy `news-digest` detritus on ana-docker; watch nh3-nas
`/volume1` (was 65%; recheck before ~80%); the `docker push 60s client-side
ceiling` mystery remains uninstrumented.
- **Granite-FP8 + observability session — all LIVE & committed (`34a43a0`, `9171e6a`).**
- **Granite 4.1 8B FP8 is the production summarizer** (`vllm-granite` :8004, ana-ml2 GPU 1,
50K ctx, CUDA graphs) — replaced phi4-mini, validated by brokkr (valid_format 1.0, FP8 stays).
- **LiteLLM gateway** (:4000) routes `granite-4.1-8b`→vLLM (explicit entry shadows the `*`
wildcard) + **Langfuse v3 wired** (ana-docker:3001, "LLM Throughput (tok/s)" dashboard built).
- **GPU-1 retuned** (trio over-provisioned KV trimmed) → granite runs with CUDA graphs + ~10 GB
free as a future Granite-text-LoRA hedge. Streaming through the gateway confirmed (TTFT 0.24s).
- **ana-docker pruned** 77 GB (unused images + build cache; disk 83%→49%) to fit ClickHouse.
- **Worldtree summarizer repoint — NO instance change now; DEFERRED to Worldtree #254** (see Recent
decisions). worldtree-dev will ping with the providers.yaml + consumer config when #254 un-holds;
infra-ops applies to the personal/demo/pinned bind mounts (vh@10.250.50.152, `/opt/worldtree*/config`).
- **Commits unpushed** (`34a43a0`, `9171e6a`, nevermore `d3e19b8` in its repo) — operator's call to push.
- **Operator flagged "new work to do"** for the next session — this snapshot is the handoff.
- **Disclosed-keys hygiene queue** (rotate at convenience): HF token `hf_HBl…` (lkraven's), `/tmp/
wt-personal-skaldsong-prod.key`, Worldtree `Z_AI_API_KEY`, Gitea runner reg token, `MINIFLUX_PASSWORD`
(leaked twice). (sk-corvid + the langfuse/vastblueai-gateway keys are dev-enclosed — leakage deprioritized.)
- **Still open from prior:** clean legacy `news-digest` on ana-docker; watch nh3-nas `/volume1`; **pin
llama-swap to GPU 0** for clean GPU-1 separation; the `docker push 60s ceiling` mystery uninstrumented.
## Recent decisions
- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer** (supersedes
the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not
Q4/AWQ** — primary workload (agent memory + summarization) is high-concurrency, where FP8-on-Ada
scales ~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a
config bug — placement/kernel/contention all ruled out). vLLM `vllm-granite` :8004 GPU 1, official
IBM compressed-tensors FP8, CUDA graphs. **GPU-1 retune** (trio utils 0.2/0.2/0.3→0.07/0.07/0.18,
granite 0.36) freed ~10 GB → CUDA graphs + a Granite-text-LoRA hedge. nevermore repointed. (`34a43a0`,
auto-memory `reference_ana_ml2_vllm_granite`)
- `[2026-06-05]` **Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI**; LiteLLM
`success_callback:[langfuse]` live (project `gateway`). Pretty prompt/completion/reasoning traces +
an `outputTokensPerSecond` tok/s dashboard. NOT a prerequisite — spend_logs already capture
tokens+latency. (`9171e6a`, auto-memory `reference_litellm_gateway`)
- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any
found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memory
`feedback_avoid_ollama`)
- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent).
FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI;
eshpfi keeps the `comfyui` stack compose, comfy-dev owns the model/workflow knowledge. (auto-memory
`reference_irv_ml1_ampere_quant`)
- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is
the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to
claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands
the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts.
**CORRECTION to the 2026-06-04 "deploys ALL CICD" line:** the bind-mount CONFIGS (providers.yaml,
vh-owned on corviduo `/opt/worldtree*/config`) ARE infra-ops's to apply directly — only the
app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`)
- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent;
granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from
Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`)
@@ -206,6 +216,31 @@ _37 older entries archived to archival-memory.md._
## Tried and abandoned
- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — it fills the KV cache to the
`--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs
AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36
with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR `--enforce-eager`
(no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores
need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`)
- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres
rows (`dashboards`/`dashboard_widgets`); build by cloning a default-dashboard row + swapping the
measure. tok/s is NOT a per-generation field (null on the observation) — it's the
`outputTokensPerSecond` MEASURE, computed at metrics-API/dashboard query time; no native per-call
tok/s display exists (streaming doesn't change that). langfuse-web needs `HOSTNAME=0.0.0.0` (Next.js
standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host
3000 is gitea's → langfuse on 3001.
- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh +
corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read
world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are
vh-owned (editable), the `.env` is deploy-owned 600 (vh can't edit it, no sudo).
- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL` on
the containers; it defaults to claude-haiku in code, opt-in via config not `.env`. Don't trust an
".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection
corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.)
- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF —
the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the
model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline