24 KiB
Persistent memory — eshpfi-management
Last updated: 2026-06-14
Repo purpose
Reference workspace for PFI infrastructure: server inventory, canonical
Docker Compose stacks, ops playbooks, and conventions. Authoritative
copies of compose files live on the servers under
/opt/docker/compose/<stack>/; this repo mirrors them for version
control, editing, planning, and CI-driven deploys.
Tools and conventions
Sister repos (separate gitea repos, deployed by playbooks here):
| Repo | Role | CI status |
|---|---|---|
vh/task-board |
MCP + web dashboard for assistant task state (port 7878) | push-to-main → CI deploys (2026-04-29) |
vh/vor |
Inquisitor UI sidecar (port 7879) | push-to-main → CI deploys (2026-04-29) |
vh/nevermore |
Twice-daily LLM-curated briefing (port 8181, replaces news-digest) | push-to-main → CI deploys (2026-04-30) |
vh/asset-engine |
Internal control plane over inference services (port 8200, LAN-direct) | push-to-main → CI deploys (2026-05-12) |
vh/althing |
Inter-agent message bus (chamber UI port 7881, forseti + agent-runner daemons, valkey IPC) | push-to-main → CI deploys (2026-05-14) |
vh/mead-hall |
Bifrost tool-provider sidecar (port 5173 on dev VM 10.100.10.50) | push-to-main → CI deploys (2026-05-16) |
vh/skaldsong |
Wizard + reader surface (port 8300, ana-docker, registry-pull pattern) | push-to-main → CI deploys (2026-05-19) |
vh/worldtree |
Conversation API (corviduo-dev demo :8080 / personal :8081 / pinned :8082) — Heimdall auth, Bifrost integration | push-to-main → CI deploys |
vh/yt-voice-clipper |
YouTube → diarized voice-clip dataset builder + audition console (irv-ml1 :8000) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (2026-06-03) — see docs/runbooks/ytvc-autodeploy.md |
vh/arbo |
Catalog-driven ComfyUI engine (irv-ml1 :8201, comfy-dev owns engine/catalog/image) | push-to-main → gitea-webhook auto-deploy to irv-ml1 (catalog-only restart; infra side = stacks/arbo/) |
(vh/volva + Heid were re-architected from systemd daemons to Claude Code
session orchestrators 2026-06-08; their nh3-dev .service units were removed —
no longer deployed sidecars here. See Recent decisions.)
-
Two-layer backups — Backrest orchestrates restic for file+DB (5 fleet repos, daily 01:00 PDT); PBS-ANA primary + PBS-NH3 DR mirror for VM images. ana-nas is the SPOF for postgres + PBS-ANA datastore + cross-site restic targets — see
docs/runbooks/disaster-recovery.mdfor the blast-radius matrix. -
pull-hf-repo.yamlis the canonical "get a HuggingFace model/dataset onto ana-ml2's shared cache at/tank/aimodels/huggingface/" playbook. Supports--var repo_type=model|dataset|space. Replaces ad-hochuggingface_hub.snapshot_downloadpatterns. -
Worldtree admin auth — per-instance. Each Worldtree deployment (demo :8080, personal :8081, pinned :8082) has its own Heimdall registry and its own bootstrap admin key. Infra-ops's stored long-lived admin key (
key_id 61419c92) atana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-adminauths against demo only. For personal-instance admin ops, fetch the bootstrap admin per-op viadocker exec worldtree-personal-worldtree-api-1 printenv WORLDTREE_BOOTSTRAP_ADMIN_KEYon corviduo-dev. Used forPOST /admin/keys, admin diagnostics (/admin/sessions/<id>/{bifrost,tools}, etc.). -
Per-project user keys against personal Worldtree (issued 2026-05-19):
skaldsong:79744637(nh3-dev iteration),skaldsong:7c1dbbbe(ana-docker prod),althing:50d85460,mead-hall:a360822d. Sameuser_id=skaldsongacross both skaldsong keys → shared Heimdall agent slot; differentkey_id→ independently rotatable. Pattern: mint via/admin/keys, drop value to/tmp/wt-personal-<name>.keymode 600, dev collects + shreds (DO NOT cat to chat transcript). -
Skaldsong CD pattern (registry-pull). Differs from althing / asset-engine which build-on-host. vh/skaldsong's CI builds and pushes
gitea.phasefinal.com/vh/skaldsong:<sha>+:latest;playbooks/deploy-skaldsong.yamlon ana-docker pulls + recreates. SHA-pin only (no:latesthealth-gated advance yet). Prereq: host needsdocker login gitea.phasefinal.comonce (read:package PAT) — not currently in the workflow. -
gitea internal route for fleet hosts. gitea is a container on ana-docker — git-SSH
10.250.50.70:222, HTTP:3000. Fleet/colo hosts must use this internal route, NOT publicgitea.phasefinal.com(38.120.12.44, ana-srv1) — the public path fail2bans the host egress IP and wedges webhook deploys.:22on10.250.50.70is ana-docker's HOST sshd, not gitea. Full gotcha indocs/orientation.md→ Git/gitea. -
docker-as-root pattern (for ops that have no admin API, e.g.
SqliteUserStore.set_bifrost_credentials): on hosts where the SSH user is in thedockergroup but lacks passwordless sudo, rundocker run --rm -v <target-dir>:/wt -v /var/run/docker.sock:/var/run/docker.sock docker:cli sh -c "..."to edit deploy-owned files without sudo. Documented with security warning inservers/corviduo-dev/README.md. docker-group membership is effectively root via bind-mount; treat as a sudo-equivalent grant. Foot-gun: when runningdocker composeinside this sandbox, any relative path in compose.yaml (e.g.${WORLDTREE_CONFIG_DIR:-./config}) resolves against the sandbox CWD, but Docker daemon interprets the resulting path against the HOST filesystem. Always pass-e VAR=/abs/pathto the docker run invocation for any relative-default config dir. -
scripts/elwaysudo handling — elway prompts for the sudo password ONCE viagetpassbefore the firstsudo: truestep. That prompt is interactive → elway can't run unattended from a non-TTY tool if any step needs sudo. For sudo-free playbooks (nosudo: truesteps) it runs fully non-interactive over key SSH. To create root-owned dirs WITHOUT host sudo, use the docker-daemon-root trick:docker run --rm -v /worktank:/mnt alpine sh -c 'mkdir -p /mnt/<x> && chown -R 1000:1000 /mnt/<x>'.
Current state / in-flight
As of 2026-06-14:
-
Arbo prod is current + the deploy loop is closed. Backend v0.11.6 / frontend v0.11.8 (frontend rides the git mount). Auth is OFF (WireGuard boundary). Pipeline is ban-immune (internal gitea route) + restart-correct (catalog-only) + version-controlled (
stacks/arbo/). OPEN: the gitea registry has never held an arbo image — every deploy is local-image-only on irv-ml1, a rollback SPOF. Operator may mint a vhpackage:writePAT to backfill; the internal-route fix made it non-blocking, so NOT done. The public-IP fail2ban ban on irv-ml1's egress (38.120.94.3) on the gitea host — operator was handling the unban, but the internal-route repoint made it MOOT; ban-cleanup is optional hygiene (status unverified). -
R16 vmoan pilot is CLOSED — v1 (default decode) is the final Chatterbox-tag inline artifact (see Recent decisions). The only forward thread: a splice-pivot de-risk yield-probe brokkr surfaced 2026-06-14 (45 standalone-NVV clips + speaker-embedding/acoustic metrics) — surfaced to operator, awaiting go; it needs Resemblyzer/ECAPA + librosa/parselmouth installs in the harness venv.
-
ana-ml2 on Blackwell (dual RTX PRO 6000 Blackwell Max-Q, 96 GB ea, cc 12.0). GPU 0 held free for large-model hot-loads (llama-swap pinned); GPU 1 steady tenants — granite 131k + qwen3.5-vl 65k + embed/rerank/reward trio, ~3.5 GB free. CLAUDE.md GPU-spec doc-fix LANDED (
355a240). -
R17 v2 corpus characterization still on the local irv-ml1 branch, push HELD (
r17-v2-characterization). No-push rule AND ASR-content-exposure call on the intimate-audio batch. Awaiting brokkr-collect or explicit push approval. (LOCAL-ONLY soundgasm-derived corpus — distribution barred, never echo transcripts to the bus.) -
Mac Pro migration is still the big open project —
migration-plan.md(repo root): WORKSTATION-ONLY move of nh3-dev's Hat-1 dev env to an M2 Ultra Mac Pro on the 10.100 subnet. Hat-2 fleet sidecars STAY on the Linux VM. Phase 0 (push pushables + confirm the ~6 local-only repos) is the only time-sensitive step. Cutover =rsyncworking trees, NOT re-clone. Everything else waits on hardware. -
sglang-vs-vLLM bench stack staged but parked (
5f049cb) — the originating SGLang-RadixAttention question resolved (vLLM v1 defaults prefix-caching ON, now pinned explicit on granite+qwen). NVFP4 chase abandoned. MAX_JOBS-on-shared-prod foot-gun caps any from-source build on ana-ml2. -
pi + GLM 5.1 harness live on nh3-dev —
glmlauncher runs pi againstglm-5.1(thinking-off) via the LiteLLM gateway;glm-5.1-reasoningfor opt-in thinking. -
Worldtree config-propagation lane is mature + humming — pre-merge delta-ping → sync-on-merge to demo+personal bind-mounts; pinned stays pre-cutover.
-
Disclosed-keys hygiene queue (rotate at convenience): HF token, the wt-personal keys, Gitea runner reg token, MINIFLUX_PASSWORD, ZAI_API_KEY. (The shared all-agents LiteLLM key is intentional, not hygiene-debt.)
-
Still open from prior: clean legacy
news-digeston ana-docker; thedocker push 60s ceilingmystery uninstrumented.
Recent decisions
-
[2026-06-14]R16 vmoan inline-generation arc CLOSED — v1 at default decode (rep1.2/temp0.8) is the final Chatterbox-tag inline artifact. Operator's ear rejected every alternative: v2/v3 windowing (omission vs coherence-loss), v4 multi-tag (cohesion held but lost to capacity-competition), emergent inline-token modulation (degenerates, not modulates), and the gen-time decode-polish sweep (soft tamers cut the NVV itself — same omission family as v2; p0 baseline beat p1). All adapters v1–v4 +tokenizer.json.v3bakpreserved onirv-ml1:~/r16-vmoan-harness. Likely-next direction (deferred, NOT formalized): generate→bin→splice + one-shot-clone NVV pipeline routing around the inline-coherence wall. Tracking: brokkr R16 journal + althing thread01KV010WGSSMPWRNCPAGSPK15Y. -
[2026-06-14]Arbo deploy pipeline fixed, hardened, and version-controlled. Prod rebuilt v0.11.1 → v0.11.6 backend; the webhook machinery (arbo-deploy.sh+arbo-webhook.py, :9009 HMAC listener) is now repo-tracked atstacks/arbo/(was host-only = recoverability foot-gun). Deploy reaches gitea via the INTERNAL route (10.250.50.70:222) and restarts the engine ONLY oncatalog/changes (graphs/frontend per-request; warn onsrc/|Dockerfileonly — pyproject/uv.lock churn every commit). Operator kept arbo stack ownership in eshpfi (not migrated to comfy-dev's repo). Secret +.envstay host-only. Tracking:6d66bc2,6e58e57,stacks/arbo/READMEQ5. -
[2026-06-13]Arbo prod bearer auth turned OFF — WireGuard is the access boundary (operator decision; reverses ADR-0001's "closed the open-auth hole"). ENGINE_TOKEN must be ABSENT, not empty (empty-string still gates) — removed from BOTH the host.envAND the composeenvironment:injection line. Original token backed up atirv-ml1:/opt/docker/compose/arbo/.env.pre-auth-off.bak; comfy-dev updated their ADR-0001. Tracking:db97899+playbooks/arbo-disable-engine-token.yaml. -
[2026-06-13]Storetank image-models archive DECOMMISSIONED; arbo is the single live ComfyUI model tree (502 G). Curated/storetank/image-models/comfy(was 919 G, the native/opt/ComfyUI/modelssymlink target) → killed everything superseded by arbo's current gen (Hunyuan, WAN2.1, FLUX.1, Chroma, SD3.5, orphaned umt5+llava ≈ 739 G) + migrated the keepers (gen-agnostic utilities + the SDXL/Pony stack, 177 G) into/storetank/arbo/models(same-fs move, skip-existing protects prod). Tracking:docs/storetank-image-models-archive.md+docs/arbo-comfyui-model-catalog.md(1902425→5007ec1). -
[2026-06-13]GRANITE_KEY provisioned to comfy-dev's nh3-dev dev env at~/.arbo_granite_key(0600) for the hero gen+judge script — verbatim copy of the prodarbo-prompt-enhancevkey (now extended to reach BOTHgranite-4.1-8bANDqwen3.5-9b-fp8); nothing minted. The vkey README's "granite-only" wording was stale → corrected (f32c6dd). -
[2026-06-13]ana-ml2 upgraded Ada → dual RTX PRO 6000 Blackwell Max-Q (96 GB each, cc 12.0 / sm_120; was dual RTX 6000 Ada 48 GB / cc 8.9 — confirmed live vianvidia-smi). Unlocks NVFP4 (FP4 tensor cores) and doubles VRAM headroom. CLAUDE.md GPU-spec doc-fix LANDED355a240(operator). Tracking:19a07b9/1e2a3a1("Blackwell 96GB"). -
[2026-06-13]NVFP4-W4A4 is infeasible for Granite — FP8 stays the Granite-on-Blackwell format. W4A4 collapses at 30k context, proven producer-independent (modelopt AND llm-compressor both clean-NONE from the same BF16 base + wikitext-2k calib). No 4-bit wins both axes: W4A4 = quality collapse; W4A16-NVFP4/AWQ = weight-only dequant → bf16 (no FP4-core speedup). 30B retired. (auto-memoryreference_nvfp4_w4a4_granite_infeasible) -
[2026-06-13]Qwen3.5-9B VL (FP8) deployed on ana-ml2 GPU 1 —qwen35-vlstack, :8007, gateway aliasqwen3.5-9b-fp8. Pinned nightly digest, not:latest: the stable release quantizes the VL vision tower under--quantization fp8→ garbage vision (LM fine); the nightly correctly excludes it. Re-pin + drop the pin once that exclusion lands stable. (2e3dcc2) -
[2026-06-13]comfyui 325 G model tree migrated worktank →/storetank/arbo(worktank 97% → 26%).arbois the consuming app; overlay bind-mount viaCOMFYUI_MODELS_DIR. (38186be) (See the 2026-06-13 archive-decommission decision above — this tree later absorbed the storetank-archive keepers, reaching 502 G.) -
[2026-06-13]GPU layout settled on the Blackwell box. GPU 0 held free for large-model hot-loads (llama-swap pinned,edf0f91); GPU 1 steady-tenant — granite 131k ctx, qwen 65k, embed/rerank/reward trio, ~3.5 GB free (1e2a3a1,19a07b9; trio re-floored for 96 GB, 20×-parallel-stable). embed/rerank left at floor — long docs chunked BEFORE embedding. max-model-len is a ceiling not a reservation. -
[2026-06-13]Prefix caching pinned explicit on granite + qwen — benched ~6.5× faster TTFT on a shared ~4.5k-token summarizer template; soft/evictable, neutral when prefixes don't repeat. vLLM v1 defaults it ON (granite) but the qwen nightly defaults OFF — pin both. (a9a2be7) -
[2026-06-13]granite-4.1-8b listed as the always-available summarizer/classifier + a shared all-agents key minted (operator-directed). Global~/.claude/CLAUDE.mdGlobal-tools entry; key aliasall-agents-local, scoped to the FREE local models only (granite + qwen-vision + embed/rerank, NOT paid GLM), internal-gateway-only, rotatable. (auto-memoryreference_litellm_gateway) -
[2026-06-13]arbo engine + frontend stack stood up (ADR-0001) — irv-ml1 co-located inference engine (ee57e69), python-based healthcheck (slim image, no curl/wget,bdb3312), frontend ro-mounted from the checkout (922e8ad, ADR-0001 D2). -
[2026-06-11]GLM thinking inverted at the LiteLLM gateway (operator call):glm-5.1defaults thinking-OFF;glm-5.1-reasoning= same z.ai upstream, thinking ON. Mechanism:litellm_params.extra_body:{thinking:{type:disabled}}—drop_paramsstrips a top-levelthinking/reasoning_effortbut forwardsextra_bodyverbatim to z.ai. Shared-gateway change. (95b2701, auto-memoryreference_litellm_gateway) -
[2026-06-11]pi coding agent installed on nh3-dev as a GLM 5.1 harness —@earendil-works/pi-coding-agentvia bun (npm's global prefix is/usr→ needs sudo, bun avoids it). Config~/.pi/agent/models.json, launcher~/.local/bin/glm. -
[2026-06-11]z.ai web-tools (regin) = z.ai hosted MCP path, NOT the/paas/v4Tool API. WORKS: MCP streamable-HTTP athttps://api.z.ai/api/mcp/{web_search_prime,web_reader}/mcp,Authorization: Bearer $ZAI_API_KEY(the MCP key, distinct fromZ_AI_API_KEYthe LLM key). Reference impl = Worldtree's Leif agent. -
[2026-06-10]Mac Pro migration framed: workstation-only (M2 Ultra ARM, racked NH3 on-subnet); sidecars stay Linux.migration-plan.md. (See in-flight.) -
[2026-06-10]Worldtree deployed-config propagation is infra-ops's OWNED lane (operator ruling). worldtree-dev pings the config delta pre-merge; infra-ops syncsconfig/*.yamlfrom MERGED canonical to the/opt/worldtree*/configbind-mounts on demo+personal. The v0.33.8 9-HOUR demo outage (amodel_roles.yamlstartup-dep that never reached the VM) is the failure mode this prevents. providers.yaml stays hand-tuned. corviduo emergency-ops =ssh vh@10.250.50.152, docker no-sudo. (auto-memoryreference_worldtree_deploys_cicd,reference_corviduo_dev_emergency_ops) -
[2026-06-09]LiteLLM scoped virtual keys issued to consumers (operator-authorized):brokkr-smithy(all-proxy-models),arbo-prompt-enhance(comfy-dev — granite, later extended to qwen-vision). Mint via/key/generate(mastersk-corvid), scope-restricted + rotatable, value → 600 file never the bus. (auto-memoryreference_litellm_gateway) -
[2026-06-08]volva.service + heid.service removed from nh3-dev — vestigial systemd daemons; Heid/Volva re-architected from Python pollers to Claude Code session orchestrators (heid12aa5a9); volva.service was crash-looping 203/EXEC. (6e2f80e) -
[2026-06-05]Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer. Beat phi4 on precision in brokkr's R15 P03. Staying FP8, not Q4/AWQ — primary workload is high-concurrency, where FP8 scales ~linearly (2010 tok/s @ C=32). vLLMvllm-granite:8004 GPU 1, official IBM compressed-tensors FP8. (Then on Ada; box has since gone Blackwell.) (34a43a0, auto-memoryreference_ana_ml2_vllm_granite) -
[2026-06-05]Langfuse v3 on ana-docker (:3001) as the gateway trace UI; LiteLLMsuccess_callback:[langfuse]live. Pretty traces + tok/s dashboard. NOT a prerequisite (spend_logs already capture tokens+latency). (9171e6a) -
[2026-06-05]Ollama BANNED fleet-wide (operator directive) — never stand one up; tear down any found; serve via llama-swap or vLLM. (auto-memoryfeedback_avoid_ollama) -
[2026-06-05]ComfyUI / FLUX.2 work split to~/development/comfy-dev(dedicated repo + agent). eshpfi keeps thecomfyui/arbostack compose; comfy-dev owns the model/workflow knowledge. (auto-memoryreference_irv_ml1_ampere_quant) -
[2026-06-05]Worldtree summarizer config refresh DEFERRED to Worldtree #254 (granite-4.1-8b is the structured-output profile, ON HOLD, no live consumer). Bind-mount CONFIGS (providers.yaml, vh-owned) ARE infra-ops's to apply directly — only the app/image DEPLOY is CICD; the.envis deploy-owned. (auto-memoryreference_worldtree_deploys_cicd)
50 older entries archived to archival-memory.md.
Tried and abandoned
-
[2026-06-14]Fleet/colo hosts must reach gitea over the INTERNAL route, NOT the public IP.gitea.phasefinal.com= public38.120.12.44(ana-srv1); gitea is a container on ana-docker, git-SSH10.250.50.70:222+ HTTP:3000. A fleet host egressing to public:22gets fail2ban-banned after any retrying git loop → silently wedges webhook auto-deploys (git fetchtimes out underset -euo pipefail, aborts before reset). Bit irv-ml1's arbo deploy.:22on10.250.50.70is ana-docker's HOST sshd (deploy key → Permission denied), NOT gitea. Documenteddocs/orientation.md(6e58e57). -
[2026-06-14]Chatterbox-Turbo decode-knob foot-guns (R16 v1-polish + emergent probes): the turbo length cap ismax_gen_len(default 1000) ont3.inference_turbo, NOTmax_new_tokens— andtts_turbo.generatedoes NOT forward it (wrap inference_turbo to cap).rep_pen 2.0 / temp 0.5BACKFIRES (degenerate 24 s run-on). Soft decode tamers cut the NVV ITSELF, not just the run-on tail (operator: "p1 trims the moaning too") — gen-time polish can't beat v1's defaults. Inline base-NVV tokens DEGENERATE (moan-cascade + gibberish), they don't modulate the surrounding words. -
[2026-06-13]Loading an old LoRA adapter after a vocab bump fails on embedding size. The harness config +tokenizer.jsonare now atnew_vocab_size=50279(v4 multi-tag); the v1/v2/v3 adapters are 50277. To load v1 (the accepted artifact), setcfg.new_vocab_size=50277beforeload_finetuned_engine_lora(else PeftModel state_dict size mismatch).tokenizer.json.v3bakis the 50277 tokenizer for a clean restore. -
[2026-06-13]Heavy from-source compile (MAX_JOBS=128) on the shared PROD GPU box PINS it — load hit 187, prod vLLM restarted, killed an in-flight quant. ana-ml2 hosts live inference; never run a big build there at full parallelism. CapMAX_JOBS≤32, build off-box, or cgroup-constrain. -
[2026-06-13]--quantization fp8on a VL model can quantize the VISION TOWER → garbage vision (Qwen3.5-VL on stable vLLM: gray-grid output; LM answers text fine, so it "looks" healthy). The nightly excludes the vision tower. Validate the VISION path on a quantized VLM, not just text — and pin the engine digest with the exclusion. -
[2026-06-13]vLLM's--gpu-memory-utilizationis checked against FREE VRAM at startup, not total — on a shared card, growing one service before trimming a co-tenant OOMs. Trim the shrinking service FIRST, then grow. Size to the FREE budget. -
[2026-06-13]Thevllm/vllm-openaientrypoint is already["vllm","serve"]— composecommand:supplies the model as the first POSITIONAL arg + flags; a secondserve/--model X→ "unrecognized arguments". Same-class:teemasks the real exit code (use>); HFdatasetsrejects barewikitext(needsSalesforce/wikitext). -
[2026-06-13]Chatterbox-Turbo LoRA finetune: the repo'ssetup.pyloads the WRONG tokenizer — pulls gpt2-medium + a grapheme merge file instead of the chatterbox-turbo GPT2 tokenizer (vocab.json+merges.txt, len 50276). Fix = override + delete the graphemetokenizer.json;[vmoan]→ new_vocab_size 50277 (1-row resize), lora_r 64 / alpha 128, modules_to_save=[text_emb,text_head]. Unique-stem corpus collision needs{index}_{stem}IDs. (irv-ml1:~/r16-vmoan-harness) -
[2026-06-11]A completion-pollwhile pgrep -f <scriptname>SELF-MATCHES its own remote shell argv — its ownpgrep -falways finds itself → the loop never exits. Use a match pattern ABSENT from the poll command (the python stage, or a sentinel file), not the driver's own name. -
[2026-06-08]Demucsuv pip install demucspulls torch 2.12/torchaudio 2.11 →ta.save()requires torchcodec → dies AFTER separating (0 stems, rc=1). Fix = pintorch==torchaudio==2.4.1+UV_LINK_MODE=copy. Validate the SAVE path, not just import + GPU inference, on a bleeding-edge torch. -
[2026-06-05]vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU — fills KV to the--gpu-memory-utilizationbudget WITHOUT reserving graph-capture memory, socapture_modelOOMs AFTER weights+KV load (crash-loops). Fix: free co-tenant room OR--enforce-eager. FP8 single-stream is batch-1 GEMV (memory-bound) → Q4 wins single-stream by physics; FP8 wins under concurrency. (reference_ana_ml2_vllm_granite) -
[2026-06-05]Langfuse has NO public dashboard-creation API — dashboards/widgets are postgres rows; clone a default + swap the measure. tok/s is theoutputTokensPerSecondMEASURE (metrics-API/dashboard query time), not a per-generation field. langfuse-web needsHOSTNAME=0.0.0.0. Host 3000 is gitea's → langfuse on 3001. -
[2026-06-05]sudoover non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD (esh + corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read world-readable files WITHOUT sudo. corviduo ssh =vh@10.250.50.152; bind-mount configs are vh-owned, the.envis deploy-owned 600. -
[2026-06-05]Worldtree summarizer-model is NOT an env var — noWORLDTREE_SUMMARIZER_MODEL; defaults to claude-haiku in code, opt-in via config not.env. Inspect the live container env + vh-owned config files first.
53 older entries archived to archival-memory.md.