27 KiB
Persistent memory — eshpfi-management
Last updated: 2026-06-13
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) | push-to-main → CI deploys (2026-05-14) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration | push-to-main → CI deploys |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. For personal-instance admin ops, fetch the bootstrap admin per-op viadocker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEYon corviduo-dev. Used forPOST /admin/keys, admin diagnostics (/admin/sessions/<id>/{bifrost,tools}, etc.). -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637(nh3-dev iteration),skaldsong:7c1dbbbe(ana-docker prod),althing:50d85460,mead-hall:a360822d. Sameuser_id=skaldsongacross both skaldsong keys → shared Heimdall agent slot; differentkey_id→ independently rotatable. Pattern: mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only (no:latesthealth-gated advance yet). Prereq: host needsdocker login gitea.phasefinal.comonce (read:package PAT) — not currently in the workflow. -
docker-as-root pattern (for ops that have no admin API, e.g.
SqliteUserStore.set_bifrost_credentials): on hosts where the SSH user is in thedockergroup but lacks passwordless sudo, rundocker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..."to edit deploy-owned files without sudo. Documented with security warning inservers/corviduo-dev/README.md. docker-group membership is effectively root via bind-mount; treat as a sudo-equivalent grant. Foot-gun: when runningdocker composeinside this sandbox, any relative path in compose.yaml (e.g.${WORLDTREE_CONFIG_DIR:-./config}) resolves against the sandbox CWD, but Docker daemon interprets the resulting path against the HOST filesystem. Always pass-e VAR=/abs/pathto the docker run invocation for any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep. That prompt is interactive → elway can't run unattended from a non-TTY tool if any step needs sudo. For sudo-free playbooks (nosudo: truesteps) it runs fully non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo, use the docker-daemon-root trick:docker run --rm -v /worktank:/mnt alpine sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'.
Current state / in-flight
As of 2026-06-13:
-
ana-ml2 went Ada → Blackwell (dual NVIDIA RTX PRO 6000 Blackwell Max-Q, 96 GB each, cc 12.0 / sm_120 — was dual RTX 6000 Ada 48 GB / cc 8.9, confirmed live via
nvidia-smi). GPU layout now: GPU 0 held free for large-model hot-loads (llama-swap pinned there,edf0f91); GPU 1 is the steady-tenant card — granite (131k ctx) + qwen3.5-vl (65k) + the embed/rerank/reward trio, ~3.5 GB free after the rebalance. CLAUDE.md's server table + the granite/qwen "Ada (cc 8.9)" compose comments are STALE — doc-fix offered, pending operator go-ahead (/snapshot doesn't auto-edit CLAUDE.md). -
R16 "vmoan" TTS LoRA POC — awaiting the operator's ear (Brokkr's designated gate). Two auditions on NFS:
/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v1.wav(42 s, run-on) vs_v2.wav(10 s, tight). v2 fixed the run-on; Brokkr holds the pilot verdict on the operator's listen. Grab:scp irv-ml1:/mnt/smithy/scratch/r16-vmoan-pilot/vmoan_audition_v{1,2}.wav .. LOCAL-ONLY corpus (soundgasm-derived) — distribution barred; never echo transcripts to the bus. -
R17 v2 corpus characterization still on the local irv-ml1 branch, push HELD (
r17-v2-character ization, 18 MULTI / 101 ELIGIBLE + 19936 candidates +speech.jsonASR). No-push rule AND ASR of the intimate-audio batch (content-exposure call). Awaiting brokkr-collect or explicit push approval. -
Mac Pro migration is still the big open project —
migration-plan.md(repo root): WORKSTATION- ONLY move of nh3-dev's Hat-1 dev env (~40 repos, all of~/.claude, dotfiles, toolchain) to an M2 Ultra Mac Pro Rack on the 10.100 subnet. Hat-2 fleet sidecars (egress SOCKS5, mead-hall, ttyd seats,althing-forseti) STAY on the Linux VM. Phase 0 (push pushables + confirm the ~6 local-only repos) is the only time-sensitive step. Cutover =rsyncworking trees, NOT re-clone (re-clone loses unsaved work in ~15 repos). macOS/home/lkraven→/Users/lkravenrepath. Everything else waits on hardware. -
sglang-vs-vLLM bench stack staged but parked (
5f049cb) — the originating question (SGLang RadixAttention caching) resolved without it: vLLM v1 already defaults prefix-caching ON, now pinned explicit on granite+qwen. The NVFP4 chase is abandoned. Bench stack stays staged if ever revisited; the MAX_JOBS-on-shared-prod foot-gun (see Tried) caps any future from-source build on ana-ml2. -
pi + GLM 5.1 harness live on nh3-dev —
glmlauncher runs pi againstglm-5.1(thinking-off) via the LiteLLM gateway;glm-5.1-reasoningfor opt-in thinking. -
Worldtree config-propagation lane is mature + humming — pre-merge delta-ping → sync-on-merge to the demo+personal bind-mounts; pinned stays pre-cutover (full migration at its next re-image).
-
Disclosed-keys hygiene queue (rotate at convenience): HF token, the wt-personal keys, Gitea runner reg token, MINIFLUX_PASSWORD, ZAI_API_KEY. (The shared all-agents LiteLLM key is intentional, not hygiene-debt — see decisions; rotatable via infra-ops only if it leaks.)
-
Still open from prior: clean legacy
news-digeston ana-docker; thedocker push 60s ceilingmystery uninstrumented. (llama-swap GPU-0 pin — DONE this session,edf0f91.)
Recent decisions
-
[2026-06-13]ana-ml2 upgraded Ada → dual RTX PRO 6000 Blackwell Max-Q (96 GB each, cc 12.0 / sm_120; was dual RTX 6000 Ada 48 GB / cc 8.9 — confirmed live vianvidia-smi). Unlocks NVFP4 (FP4 tensor cores) and doubles VRAM headroom. CLAUDE.md's server table + the granite/qwen compose "Ada (cc 8.9)" comments are now stale; doc-fix offered, pending operator go-ahead (snapshot doesn't auto-edit CLAUDE.md). Tracking surface: this snapshot + commits19a07b9/1e2a3a1("Blackwell 96GB"). -
[2026-06-13]NVFP4-W4A4 is infeasible for Granite — FP8 stays the Granite-on-Blackwell format. W4A4 (4-bit weights + activations) collapses at 30k context, proven producer-independent (modelopt AND llm-compressor both clean-NONE from the same BF16 base + wikitext-2k calib). No 4-bit wins both axes: W4A4 = quality collapse; W4A16-NVFP4/AWQ = weight-only dequant → bf16 (no FP4-core speedup), so the tempting ~40% uplift doesn't materialize without quality loss. 30B retired (not enough quality for the VRAM/perf hit for our use case). (auto-memoryreference_nvfp4_w4a4_granite_infeasible) -
[2026-06-13]Qwen3.5-9B VL (FP8) deployed on ana-ml2 GPU 1 —qwen35-vlstack, :8007, gateway aliasqwen3.5-9b-fp8. Clean FP8 (no NVFP4 for vision). Pinned nightly digest, not:latest: the stable release quantizes the Qwen3.5-VL vision tower under--quantization fp8→ garbage vision (LM fine, "sees" noise); the nightly correctly excludes the vision tower. Re-pin to:latest+ drop the pin once that exclusion lands stable. (2e3dcc2) -
[2026-06-13]comfyui 325 G model tree migrated worktank →/storetank/arbo(worktank 97% → 26%).arbois the consuming app (ComfyUI-backed prompt/image pipeline); models live under its name on the bigger pool. Overlay bind-mount viaCOMFYUI_MODELS_DIRin the comfyui compose; inventory atdocs/arbo-comfyui-model-catalog.md(retain-large-portion-for-future-use decision). (38186be) -
[2026-06-13]GPU layout settled on the Blackwell box. GPU 0 held free for large-model hot-loads (llama-swap pinned there,edf0f91); GPU 1 is the steady-tenant card — granite 131k ctx (was 51k), qwen 65k, embed/rerank/reward trio, ~3.5 GB free after rebalance (1e2a3a1,19a07b9; trio re-floored for 96 GB, 20×-parallel-stable). embed/rerank left at floor — long docs are chunked BEFORE embedding, so a longer embed ctx buys nothing. PagedAttention note: max-model-len is a ceiling, not a reservation, so small requests aren't blocked by the big ceiling; concurrency = KV-pool-tokens / actual-request-size. -
[2026-06-13]Prefix caching pinned explicit on granite + qwen — benched ~6.5× faster TTFT (45 ms vs 292 ms) on a shared ~4.5k-token summarizer template; soft/evictable KV, neutral when prefixes don't repeat. vLLM v1:latestdefaults it ON (granite) but the qwen nightly defaults it OFF — pin both so a version flip can't silently disable it. (a9a2be7) -
[2026-06-13]granite-4.1-8b listed as the always-available summarizer/classifier endpoint + a shared all-agents key minted (operator-directed). Added to the global~/.claude/CLAUDE.mdGlobal- tools section; key aliasall-agents-local, scoped to the FREE local models only (granite + qwen- vision + embed/rerank, NOT the paid GLM), internal-gateway-only, rotatable. The standing "reach for this before spending premium tokens on low-caliber high-volume work" lever. (auto-memoryreference_litellm_gateway) -
[2026-06-13]arbo engine + frontend stack stood up (ADR-0001) — irv-ml1 co-located inference engine (ee57e69), python-based healthcheck (the slim image ships no curl/wget,bdb3312), frontend ro-mounted from the v0.11.2 checkout (922e8ad, ADR-0001 D2). arbo = the ComfyUI-consuming app whose models now live at/storetank/arbo. -
[2026-06-11]GLM thinking inverted at the LiteLLM gateway (operator call):glm-5.1defaults thinking-OFF; newglm-5.1-reasoningalias = same z.ai upstream with thinking ON (opt-in). Mechanism:litellm_params.extra_body:{thinking:{type:disabled}}— LiteLLMdrop_paramsSTRIPS a top-levelthinking/reasoning_effort, but forwardsextra_bodyverbatim to z.ai (the only channel that works; verified reasoning_tokens 0 vs >0). Shared-gateway change — affects ALL glm-5.1 callers (brokkr's all-models key included). (95b2701, auto-memoryreference_litellm_gateway) -
[2026-06-11]pi coding agent installed on nh3-dev as a GLM 5.1 harness —@earendil-works/ pi-coding-agentvia bun (user-level; npm's global prefix is/usr→ needs sudo, bun avoids it). Config~/.pi/agent/models.json(litellm provider → gateway), launcher~/.local/bin/glmsources the gateway key + selects the model. pi is OpenAI-compatible; proxy-safe compat flags for the GLM path. -
[2026-06-11]z.ai web-tools (regin) = z.ai hosted MCP path, NOT the/paas/v4Tool API. Directapi.z.ai/api/paas/v4/web_search→ 429/1113 "insufficient balance" (coding-plan keys bill tools on a separate quota path);open.bigmodel.cnis the China platform (404, different account). WORKS: MCP streamable-HTTP athttps://api.z.ai/api/mcp/{web_search_prime,web_reader}/mcp,Authorization: Bearer $ZAI_API_KEY(the MCP key — distinct fromZ_AI_API_KEYthe LLM key). Reference impl = Worldtree's Leif agent (tools/web/zai_client.py+core/clients/mcp.py). -
[2026-06-10]Mac Pro migration framed: workstation-only (M2 Ultra ARM, racked NH3 on-subnet); sidecars stay Linux.migration-plan.md. (See in-flight.) -
[2026-06-10]Worldtree deployed-config propagation is infra-ops's OWNED lane (operator ruling). worldtree-dev pings the config delta on every config-touching commit (pre-merge); infra-ops syncsconfig/*.yamlfrom MERGED canonical to the/opt/worldtree*/configbind-mounts on demo+personal (Heimdall hot-reloads policies; defaults are code-defaulted). The v0.33.8 9-HOUR demo outage — amodel_roles.yamlhard-startup-dep that shipped in canonical 3 releases earlier but never reached the VM — is the failure mode this lane prevents. providers.yaml stays hand-tuned (artemis graft on personal). corviduo-dev emergency-ops =ssh vh@10.250.50.152(alias doesn't resolve, lkraven denied, infra-ops key excluded), docker no-sudo. (auto-memoryreference_worldtree_deploys_cicd,reference_corviduo_dev_emergency_ops) -
[2026-06-09]LiteLLM scoped virtual keys issued to consumers (operator-authorized):brokkr- smithy(all-proxy-models),arbo-prompt-enhance(comfy-dev, granite-only). Pattern: mint via/key/generate(mastersk-corvid), scope-restricted + rotatable, value → 600 file never the bus. (auto-memoryreference_litellm_gateway) -
[2026-06-08]volva.service + heid.service removed from nh3-dev — vestigial systemd daemons; Heid/Volva re-architected from Python pollers to Claude Code session orchestrators (heid commit12aa5a9); binaries+dirs gone, volva.service was crash-looping 203/EXEC. Cleanup at heid's request (the bus-content-driven sudo got the harness guardrail; operator green-lit). (6e2f80e) -
[2026-06-05]Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer (supersedes the 2026-06-04 phi4 decision below). Beat phi4 on precision in brokkr's R15 P03. Staying FP8, not Q4/AWQ — primary workload (agent memory + summarization) is high-concurrency, where FP8 scales ~linearly (profiled 2010 tok/s @ C=32; single-stream 67.5 is batch-1 GEMV physics, not a config bug). vLLMvllm-granite:8004 GPU 1, official IBM compressed-tensors FP8, CUDA graphs. (Then on Ada cc 8.9; the box has since gone Blackwell — see 2026-06-13.) (34a43a0, auto-memoryreference_ana_ml2_vllm_granite) -
[2026-06-05]Langfuse v3 stood up on ana-docker (:3001) as the gateway trace UI; LiteLLMsuccess_callback:[langfuse]live (projectgateway). Pretty prompt/completion/reasoning traces + anoutputTokensPerSecondtok/s dashboard. NOT a prerequisite — spend_logs already capture tokens+latency. (9171e6a, auto-memoryreference_litellm_gateway) -
[2026-06-05]Ollama BANNED fleet-wide (operator directive) — never stand one up; tear down any found; serve via llama-swap or vLLM. Torn down irv-ml1 :11434 (freed 19 GB). (auto-memoryfeedback_avoid_ollama) -
[2026-06-05]ComfyUI / FLUX.2 work split to~/development/comfy-dev(dedicated repo + agent). FLUX.2-klein (fp8 + q8 GGUF, stock + uncensored encoders) installed on the irv-ml1 Docker ComfyUI; eshpfi keeps thecomfyuistack compose, comfy-dev owns the model/workflow knowledge. (auto-memoryreference_irv_ml1_ampere_quant) -
[2026-06-05]Worldtree summarizer config refresh DEFERRED to Worldtree #254 (granite-4.1-8b is the structured-output profile, ON HOLD, no live consumer; the conversation summarizer defaults to claude-haiku — the "phi4 erroring" premise was wrong). No instance changes now; worldtree-dev hands the exact providers.yaml + consumer config when #254 un-holds, infra-ops applies to the bind mounts. CORRECTION to the 2026-06-04 "deploys ALL CICD" line: the bind-mount CONFIGS (providers.yaml, vh-owned on corviduo/opt/worldtree*/config) ARE infra-ops's to apply directly — only the app/image DEPLOY is CICD; the.envis deploy-owned. (auto-memoryreference_worldtree_deploys_cicd) -
[2026-06-04]phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer/dreaming agent; granite-4-small retired from llama-swap (config-only; GGUFs on disk). 50K ctx (dropped from Phi-4's 128K max to fit GPU 1's ~10 GB free) + FP8 KV. (40a374b) -
[2026-06-04]phi4 ships the CANONICAL/official Phi-4 chat template, NOT Ollama's. Ollama's bundled template omits the system<|end|>— that flattered brokkr's R15 eval but is the DIVERGENT scaffold (Dvalin: the system<|end|>is Microsoft's intended format). Applied an Ollama-matching override then reverted — ship correct, not the benchmark quirk. (90e08f0→27eb537; "headgun" lesson in Tried.) -
[2026-06-04]infra-ops NOPASSWD-sudo identity commissioned, scoped to PFI boxes (+esh-docker-vm by operator override) — so infra-ops completes DevOps end-to-end vs handing the operator sudo steps. Dedicated key, sudo log_output, key-gated. (8c32a05) -
[2026-06-04]Worldtree demo/pinned/personal deploys are ALL CI/CD, not infra-ops — a "deploy vX.Y.Z" request to infra-ops is MISROUTED → point them back to their pipeline. The granite→phi4 repoint: worldtree-dev self-served via their CI/CD (v0.30.10). (d8d776c) -
[2026-06-04]ollama upgraded 0.9.0→0.30.4 on irv-ml1 (Ministral-3 is a Dec-2025 model the old engine refused); A6000 pinned by UUID not index (native fastest-first ≠ nvidia-smi PCI). -
[2026-06-04]brokkruser (no-sudo) on irv-ml1; R14/R15/R16 substrate moved to /home/brokkr. Persistent box services there need SYSTEM systemd units (see Tried).
44 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-06-13]Heavy from-source compile (MAX_JOBS=128sglang fork build) on the shared PROD GPU box PINS it — load hit 187, prod vLLM services restarted, killed an in-flight quant. ana-ml2 hosts live inference; never run a big build there at full parallelism. CapMAX_JOBS≤32, build off-box, or cgroup-constrain. (Operator ran the kill; infra-ops NOPASSWD-sudo confirmed working on ana-ml2 — retry infra-ops on an ssh-255 before concluding "no access.") -
[2026-06-13]--quantization fp8on a VL model can quantize the VISION TOWER → garbage vision (Qwen3.5-VL on the stable vLLM: gray-grid output; the LM answers text fine, so it "looks" healthy until you actually feed it an image). The nightly excludes the vision tower. Lesson: validate the VISION path on a quantized VLM, not just text — and pin the engine digest that has the exclusion. -
[2026-06-13]vLLM's--gpu-memory-utilizationis checked against FREE VRAM at startup (free >= util*total), not total — on a shared card, growing one service before trimming a co-tenant OOMs ("free 48.56 < desired 80.72" at util 0.85 on a half-occupied 96 GB card). Start-order matters: trim the shrinking service FIRST, then grow the other. Size to the FREE budget, not the total. -
[2026-06-13]Thevllm/vllm-openaientrypoint is already["vllm","serve"]— the composecommand:supplies the model as the first POSITIONAL arg + flags; a secondserve(or--model X) yields "unrecognized arguments". Same-class gotchas this session:teemasks the real exit code (use a>redirect to keep rc); HFdatasetsrejects barewikitext(needsSalesforce/wikitext). -
[2026-06-13]Chatterbox-Turbo LoRA finetune: the repo'ssetup.pyloads the WRONG tokenizer —merge_and_save_turbo_tokenizer()pulls gpt2-medium + a grapheme merge file (len mismatch) instead of the chatterbox-turbo GPT2 tokenizer (vocab.json+merges.txt, len 50276). Fix = override with the correct tokenizer + delete the graphemetokenizer.json;[vmoan]token → new_vocab_size 50277 (1-row resize), lora_r 64 / alpha 128, modules_to_save=[text_emb,text_head]. Also: a unique-stem corpus collision (53 rows, 32 wavs) needs{index}_{stem}IDs. (irv-ml1~/r16-vmoan-harness) -
[2026-06-11]A completion-pollwhile pgrep -f <scriptname>SELF-MATCHES its own remote shell argv — the poll's command line contains the script name, so its ownpgrep -falways finds itself → the loop never exits, the poll never fires. Use a match pattern ABSENT from the poll command (pgrep the python stage, or a sentinel file), not the driver's own name. (Caught only because the operator asked "check status"; the job had already finished cleanly.) -
[2026-06-08]Demucsuv pip install demucspulls torch 2.12/torchaudio 2.11 →ta.save()requires torchcodec → dies AFTER separating (0 stems written, rc=1). Same class as the torch-2.12 torchcodec foot-gun. Fix = pintorch==torchaudio==2.4.1(pre-torchcodec save backend) +UV_LINK_MODE=copyfor the EPERM-hardlink quirk. Lesson restated: validate the SAVE path, not just import + GPU inference, on a bleeding-edge torch. -
[2026-06-05]vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU — it fills the KV cache to the--gpu-memory-utilizationbudget WITHOUT reserving graph-capture memory, socapture_modelOOMs AFTER weights+KV load (model/KV log looks healthy, then crash-loops; saw 11 restarts at util 0.36 with 237 MB free). Fix: free co-tenant room (right-size the other vLLM services) OR--enforce-eager(no graphs, ~15-25% slower decode). FP8 single-stream is batch-1 GEMV (memory-bound, FP8 tensor cores need batch>1) → Q4 wins single-stream by physics; FP8 wins under concurrency. (reference_ana_ml2_vllm_granite) -
[2026-06-05]Langfuse has NO public dashboard-creation API — dashboards/widgets are postgres rows (dashboards/dashboard_widgets); build by cloning a default-dashboard row + swapping the measure. tok/s is NOT a per-generation field (null on the observation) — it's theoutputTokensPerSecondMEASURE, computed at metrics-API/dashboard query time; no native per-call tok/s display exists (streaming doesn't change that). langfuse-web needsHOSTNAME=0.0.0.0(Next.js standalone binds one net-IP otherwise, unreachable via the published port once also on tnet). Host 3000 is gitea's → langfuse on 3001. -
[2026-06-05]sudoover non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD (esh + corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read world-readable files WITHOUT sudo. corviduo ssh =vh@10.250.50.152; bind-mount configs are vh-owned (editable), the.envis deploy-owned 600 (vh can't edit it, no sudo). -
[2026-06-05]Worldtree summarizer-model is NOT an env var — noWORLDTREE_SUMMARIZER_MODELon the containers; it defaults to claude-haiku in code, opt-in via config not.env. Don't trust an ".env-flip" recipe — inspect the live container env + the vh-owned config files first. (Inspection corrected a wrong "summarizer erroring on phi4" premise → saved churning 3 live instances.) -
[2026-06-04]Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF — the "headgun" lesson. Ollama's phi4 template drops the system<|end|>; serving vLLM with the model's HF tokenizer template (canonical, has it) regressed brokkr's Ollama-measured R15 baseline -33pp type-F1 while valid_format held 1.0. An Ollama-matching--chat-template"fixed" it but was the WRONG fix (the bundled scaffold is the divergent one). PRINCIPLE: serve each model's canonicaltokenizer.apply_chat_template, not the bundled template — bundled ones corrupt baselines. Verify the applied prompt via vLLM/tokenize→/detokenize. (90e08f0/27eb537) -
[2026-06-04]GPU pin by INDEX is ambiguous on irv-ml1 — native CUDA orders fastest-first (A6000=0) but nvidia-smi/docker use PCI order (A6000=1), so an index pin can land on the wrong card. Pin by UUID (CUDA_VISIBLE_DEVICES=GPU-…); verify via nvidia-smi compute-apps. Check loaded-model VRAM withollama ps(Ministral-3 @ its 256K default ctx = ~30 GB; cap num_ctx). -
[2026-06-04]Persistent services on irv-ml1 need SYSTEM systemd units — the box reaps user-session processes on ssh disconnect, and--usersystemd isn't reachable over non-login ssh, so nohup/setsid/screen -dmS/systemd-run --userall die (even withenable-linger). Use/etc/systemd/system/. -
[2026-06-04]pyworld needssetuptools<81(imports the removedpkg_resources); and R/soundgen-lgfortranfails on irv-ml1 because the defaultgccis gcc-11 but only gfortran-12 is present (libgfortran.so lives only in the gcc-12 dir) → installlibgfortran-11-dev. -
[2026-06-04]homepage "crash" ≠ always NFS — a wedged container in unkillable D-state ("tried to kill container, but did not receive an exit event") can come from deadsiteMonitorwidget targets (retired ESH firewall IPs) hanging the node event loop intoexit_mmap, needing a host reboot. Check homepage's siteMonitors against retired hosts. (incident_esh_docker_nfs_boot_race)
48 older entries archived to archival-memory.md.