memory: snapshot — 2026-06-05 granite-FP8 cutover + Langfuse observability + worldtree #254 deferral
This commit is contained in:
+76
-41
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-06-04_
|
||||
_Last updated: 2026-06-05_
|
||||
|
||||
## Repo purpose
|
||||
|
||||
@@ -88,50 +88,60 @@ Sister repos (separate gitea repos, deployed by playbooks here):
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-06-04:_
|
||||
_As of 2026-06-05:_
|
||||
|
||||
- **INFRA SESSION 2026-06-04 — phi4 on vLLM, infra-ops sudo identity, R15/R16, brokkr svc.**
|
||||
Major threads (detail in cited auto-memories + commits):
|
||||
- **phi4-mini FP8 LIVE on ana-ml2 vLLM** (`vllm-phi4`, :8004, GPU 1, 50K ctx, FP8 + FP8-KV)
|
||||
as the **nevermore** summarizer/dreaming agent — superseded **granite-4-small** (removed
|
||||
from llama-swap config; GGUFs kept on disk). Template = **CANONICAL/official Phi-4** (final,
|
||||
after an Ollama-matching override applied `90e08f0` then reverted `27eb537`). nevermore
|
||||
repointed (LLAMA_SWAP_URL→:8004, MODEL→phi4-mini). auto-memory `reference_ana_ml2_vllm_phi4`.
|
||||
OPEN: brokkr re-baselining R15 P02 under canonical (eval re-regresses ~0.78→~0.45-0.57; not
|
||||
live — worldtree #252 is baseline-first); **GPU 1 tight (~10 GB free — pin llama-swap to
|
||||
GPU 0 as follow-up)**; Worldtree Vili #253 (hardcoded :9292 granite fallback, their fix).
|
||||
- **infra-ops NOPASSWD-sudo identity commissioned** across PFI boxes — `ssh infra-ops@<host>`,
|
||||
key `~/.ssh/infra-ops_ed25519`. Bootstrap `playbooks/bootstrap-infra-ops-user.yaml` +
|
||||
`scripts/bootstrap-infra-ops-fleet.sh` (tiers 1+2 live: irv-ml1/ana-ml2/ana-docker/nh3-docker/
|
||||
ana-nas + 4 app VMs; **esh-docker-vm added by operator override**). Excludes sf-*/corviduo/
|
||||
Synology. auto-memory `reference_infra_ops_sudo_identity` (`8c32a05`).
|
||||
- **R15/R16 brokkr-smithy harness stood up on irv-ml1** — ollama upgraded **0.9.0→0.30.4**
|
||||
(Ministral-3 needs it), A6000 **UUID-pinned**; R15 ollama arms (granite4.1:3b/8b, qwen3:4b,
|
||||
phi4-mini:3.8b, hf SmolLM3-GGUF, ministral-3:3b-instruct, nuextract:3.8b) + R16 R/soundgen +
|
||||
pyworld venv. `playbooks/irv-ml1-r15-r16-{nosudo,sudo}.yaml`. auto-memory `reference_irv_ml1_gpu_r14`.
|
||||
- **`brokkr` user created on irv-ml1** (no-sudo) + 24 GB R14/R15/R16 substrate migrated out of
|
||||
lkraven's home → `/home/brokkr/`; gitea pull = read-only deploy key. **brokkr-audition.service**
|
||||
(SYSTEM systemd unit, :8137) serves brokkr's R16 NVV audition UI (the morph set being auditioned).
|
||||
- **homepage incident (esh-docker-vm)** — wedged on dead siteMonitor IP (retired ESH firewall
|
||||
10.0.250.1) into unkillable D-state; host reboot cleared it; ESH-Firewall widget removed from
|
||||
`services.yaml`. (`incident_esh_docker_nfs_boot_race` updated.)
|
||||
- **Observability roadmap** `docs/roadmap.md` — Langfuse (full req/resp tracing) +
|
||||
Prometheus/Grafana off vLLM `/metrics`. Deferred; Langfuse first.
|
||||
|
||||
- **Disclosed-keys hygiene queue** — rotate at convenience: HF token `hf_HBl…`
|
||||
(lkraven's HF account) leaked into BuildKit logs during the CSM build attempt
|
||||
(logs shredded, never committed — low urgency);
|
||||
`/tmp/wt-personal-skaldsong-prod.key` on nh3-dev; mead-hall's prior Worldtree
|
||||
bearer (superseded by `a360822d`); Worldtree `Z_AI_API_KEY`/`ZAI_API_KEY`;
|
||||
chamber `forseti`/`agent_runner` api_keys (superseded by `50d85460`); Gitea
|
||||
runner registration token (`a1135753…`).
|
||||
- **Still open from prior sessions:** rotate `MINIFLUX_PASSWORD` (leaked
|
||||
twice); clean up legacy `news-digest` detritus on ana-docker; watch nh3-nas
|
||||
`/volume1` (was 65%; recheck before ~80%); the `docker push 60s client-side
|
||||
ceiling` mystery remains uninstrumented.
|
||||
- **Granite-FP8 + observability session — all LIVE & committed (`34a43a0`, `9171e6a`).**
|
||||
- **Granite 4.1 8B FP8 is the production summarizer** (`vllm-granite` :8004, ana-ml2 GPU 1,
|
||||
50K ctx, CUDA graphs) — replaced phi4-mini, validated by brokkr (valid_format 1.0, FP8 stays).
|
||||
- **LiteLLM gateway** (:4000) routes `granite-4.1-8b`→vLLM (explicit entry shadows the `*`
|
||||
wildcard) + **Langfuse v3 wired** (ana-docker:3001, "LLM Throughput (tok/s)" dashboard built).
|
||||
- **GPU-1 retuned** (trio over-provisioned KV trimmed) → granite runs with CUDA graphs + ~10 GB
|
||||
free as a future Granite-text-LoRA hedge. Streaming through the gateway confirmed (TTFT 0.24s).
|
||||
- **ana-docker pruned** 77 GB (unused images + build cache; disk 83%→49%) to fit ClickHouse.
|
||||
- **Worldtree summarizer repoint — NO instance change now; DEFERRED to Worldtree #254** (see Recent
|
||||
decisions). worldtree-dev will ping with the providers.yaml + consumer config when #254 un-holds;
|
||||
infra-ops applies to the personal/demo/pinned bind mounts (vh@10.250.50.152, `/opt/worldtree*/config`).
|
||||
- **Commits unpushed** (`34a43a0`, `9171e6a`, nevermore `d3e19b8` in its repo) — operator's call to push.
|
||||
- **Operator flagged "new work to do"** for the next session — this snapshot is the handoff.
|
||||
- **Disclosed-keys hygiene queue** (rotate at convenience): HF token `hf_HBl…` (lkraven's), `/tmp/
|
||||
wt-personal-skaldsong-prod.key`, Worldtree `Z_AI_API_KEY`, Gitea runner reg token, `MINIFLUX_PASSWORD`
|
||||
(leaked twice). (sk-corvid + the langfuse/vastblueai-gateway keys are dev-enclosed — leakage deprioritized.)
|
||||
- **Still open from prior:** clean legacy `news-digest` on ana-docker; watch nh3-nas `/volume1`; **pin
|
||||
llama-swap to GPU 0** for clean GPU-1 separation; the `docker push 60s ceiling` mystery uninstrumented.
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer** (supersedes
|
||||
the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not
|
||||
Q4/AWQ** — primary workload (agent memory + summarization) is high-concurrency, where FP8-on-Ada
|
||||
scales ~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a
|
||||
config bug — placement/kernel/contention all ruled out). vLLM `vllm-granite` :8004 GPU 1, official
|
||||
IBM compressed-tensors FP8, CUDA graphs. **GPU-1 retune** (trio utils 0.2/0.2/0.3→0.07/0.07/0.18,
|
||||
granite 0.36) freed ~10 GB → CUDA graphs + a Granite-text-LoRA hedge. nevermore repointed. (`34a43a0`,
|
||||
auto-memory `reference_ana_ml2_vllm_granite`)
|
||||
|
||||
- `[2026-06-05]` **Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI**; LiteLLM
|
||||
`success_callback:[langfuse]` live (project `gateway`). Pretty prompt/completion/reasoning traces +
|
||||
an `outputTokensPerSecond` tok/s dashboard. NOT a prerequisite — spend_logs already capture
|
||||
tokens+latency. (`9171e6a`, auto-memory `reference_litellm_gateway`)
|
||||
|
||||
- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any
|
||||
found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memory
|
||||
`feedback_avoid_ollama`)
|
||||
|
||||
- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent).
|
||||
FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI;
|
||||
eshpfi keeps the `comfyui` stack compose, comfy-dev owns the model/workflow knowledge. (auto-memory
|
||||
`reference_irv_ml1_ampere_quant`)
|
||||
|
||||
- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is
|
||||
the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to
|
||||
claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands
|
||||
the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts.
|
||||
**CORRECTION to the 2026-06-04 "deploys ALL CICD" line:** the bind-mount CONFIGS (providers.yaml,
|
||||
vh-owned on corviduo `/opt/worldtree*/config`) ARE infra-ops's to apply directly — only the
|
||||
app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`)
|
||||
|
||||
- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent;
|
||||
granite-4-small retired** from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from
|
||||
Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (`40a374b`)
|
||||
@@ -206,6 +216,31 @@ _37 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — it fills the KV cache to the
|
||||
`--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs
|
||||
AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36
|
||||
with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR `--enforce-eager`
|
||||
(no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores
|
||||
need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`)
|
||||
|
||||
- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres
|
||||
rows (`dashboards`/`dashboard_widgets`); build by cloning a default-dashboard row + swapping the
|
||||
measure. tok/s is NOT a per-generation field (null on the observation) — it's the
|
||||
`outputTokensPerSecond` MEASURE, computed at metrics-API/dashboard query time; no native per-call
|
||||
tok/s display exists (streaming doesn't change that). langfuse-web needs `HOSTNAME=0.0.0.0` (Next.js
|
||||
standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host
|
||||
3000 is gitea's → langfuse on 3001.
|
||||
|
||||
- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh +
|
||||
corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read
|
||||
world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are
|
||||
vh-owned (editable), the `.env` is deploy-owned 600 (vh can't edit it, no sudo).
|
||||
|
||||
- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL` on
|
||||
the containers; it defaults to claude-haiku in code, opt-in via config not `.env`. Don't trust an
|
||||
".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection
|
||||
corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.)
|
||||
|
||||
- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF —
|
||||
the "headgun" lesson.** Ollama's phi4 template drops the system `<|end|>`; serving vLLM with the
|
||||
model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline
|
||||
|
||||
Reference in New Issue
Block a user