Commit Graph

  • 3a08abd60d fix(lora-worker): default network_alpha to dim/2 to match proven Sindra runs vh 2026-07-06 18:38:28 -07:00
  • 888ba6a714 feat(lora-worker): stand up in-arbo LoRA training worker on irv-ml1 (arbo Phase 1 §4.1) vh 2026-07-06 18:28:04 -07:00
  • 5c64d31094 memory: snapshot — AEON Qwen3.6-27B is now gen (qwopus displaced), litbench torn down/comfyui restored vh 2026-07-06 00:57:01 -07:00
  • e6ab51c74a feat(aeon): deploy Qwen3.6-27B AEON as gen + char-rp, displace qwopus vh 2026-07-06 00:53:38 -07:00
  • 993decf3eb memory: snapshot — 2026-07-05 (T1 train venue = cloud-rec/operator-chose-ana-ml2-smoke; LitBench-RM up + comfyui displaced; character-rp shipped + #344; althing v2 herald/receiver systemd + PATH fix; glm-5.2 1M/128K; LiteLLM shared-param-mutation footgun; condensed R22 + several Recent-decisions entries) vh 2026-07-05 16:14:37 -07:00
  • 624a07e9c2 docs(litellm): record glm-5.2 canonical limits in gateway config comment vh 2026-07-05 09:09:05 -07:00
  • 3c966b2631 memory: track low-pri cleanup of inert mood.decay_rate/stale_hours from deployed WT bind-mounts vh 2026-07-03 22:14:58 -07:00
  • 3a627c6c26 memory: operator decided DEFER granite efficacy to the T1 run (no intermediate spike) vh 2026-07-02 10:33:15 -07:00
  • 7fdda2de53 memory: granite spike = mechanical-green ONLY, efficacy not validated by design (mtf-dev confirm) vh 2026-07-02 10:25:46 -07:00
  • 5b52673b75 memory: snapshot — 2026-07-02 (Deckard trial→revert to qwopus; MTP concurrency verdict = not-kept-on-shared-gen; worldtree #332 diagnosis + scoped-view/tunnel + CI-race lesson; mtf-dev granite harness spike; /books ESH mount) vh 2026-07-02 08:26:35 -07:00
  • 681eb705a2 Revert "ops(litellm): repoint gen/qwen-large/summarizer-large aliases -> qwen3.6-40b-deckard (Deckard trial)" vh 2026-07-01 08:19:58 -07:00
  • b63c48b19b ops(litellm): repoint gen/qwen-large/summarizer-large aliases -> qwen3.6-40b-deckard (Deckard trial) vh 2026-07-01 00:18:40 -07:00
  • 809c51e095 docs(pfi): sync recommended-model-settings KB to deployed gateway defaults vh 2026-06-27 09:45:10 -07:00
  • b9dcbc199f litellm(granite): revert temperature 0.1 -> 0 (research-dictated) vh 2026-06-27 09:40:47 -07:00
  • 9a772ec1ae litellm(granite): temperature 0 -> 0.1 (near-greedy floor) vh 2026-06-27 09:31:15 -07:00
  • 95d0b38d0a litellm: canonical defaults for granite + GLM (completes fleet sweep) vh 2026-06-27 08:33:43 -07:00
  • 4f094fa653 litellm: canonical general-use sampling defaults across gateway models vh 2026-06-27 08:26:46 -07:00
  • 52d5f66216 litellm(gen/qwen3.5-122-a10b): presence_penalty=1.0 anti-repetition default vh 2026-06-27 08:09:53 -07:00
  • 3239b0a613 comfyui(irv-ml1): add PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True vh 2026-06-25 07:30:12 -07:00
  • 30c883be5d memory: snapshot — 2026-06-25 (althing v0.17 / nh3-extdev model-B mesh + zellij web pilot + Worldtree #314/#322/#317 config arc; rewrite stale althing Tools-row to v0.17; archive 10 pre-session 2026-06-20 entries) vh 2026-06-25 00:14:02 -07:00
  • 13bfa4a621 memory: snapshot — 2026-06-21 (compress in-flight to current; archive 21 pre-session entries to archival-memory.md) vh 2026-06-21 00:07:57 -07:00
  • db2953e690 memory: R22 Phase B CANCELLED (Worldtree model-agnostic → no deploy path); gateway-only vh 2026-06-20 22:04:37 -07:00
  • 245b217372 memory: R22 key full-open confirmed by operator (settled) vh 2026-06-20 21:59:40 -07:00
  • 3cb54efd59 memory: R22 key re-minted persistent (stateless consumer); old orphan revoked vh 2026-06-20 21:55:00 -07:00
  • b33049ce3e memory: R22 operator steer — research gated on pragmatic/deployable outcome, not art vh 2026-06-20 21:52:17 -07:00
  • 9d65339fb2 memory: R22 stood down to gateway, full-access key minted, Phase B parked; phantom qwen3.6 entry to clean vh 2026-06-20 21:49:48 -07:00
  • d9ebe8d0f4 memory: old claude-bot token (id 15) revoked — worldtree-dev fully self-serve on one token vh 2026-06-20 17:27:47 -07:00
  • 6430c01dad memory: claude-bot issue-scope token (id 16) minted for worldtree-dev self-serve vh 2026-06-20 17:25:54 -07:00
  • ec671e86c5 memory: cb2a79a readonly-admin allow-rules re-staged to demo+personal (PDP is rule-based) vh 2026-06-20 16:43:44 -07:00
  • b13f66aea9 memory: ratatoskr flipped to readonly — re-staged 439bebf policies.yaml to personal + safe restart reload vh 2026-06-20 16:29:15 -07:00
  • 2e993ac3df memory: ratatoskr resolved (operator chose admin; worldtree-dev self-served key 90db1fbd) vh 2026-06-20 16:24:21 -07:00
  • 5a3a75b73d docs(backups): harden + live-activate /mnt/compose automount on ana-docker vh 2026-06-20 16:21:05 -07:00
  • 76b317ce3e feat(backups): freshness check + daily alert timer; record rest-server-ana recovery, fstab hardening, esh-pve-nas gap, worldtree admin-key provisioning vh 2026-06-20 16:13:13 -07:00
  • a7b4a82dec docs(backups): add backup architecture + freshness runbook; record rest-server-ana recovery + correct ana-docker sudo path vh 2026-06-20 15:56:41 -07:00
  • 8f15f6bb0d memory: capture 2026-06-20 session — Worldtree v0.37.7 demo fix, gitea notifier recovery, backup diagnosis vh 2026-06-20 15:46:14 -07:00
  • 58ec80d58a memory: snapshot — 2026-06-20 vh 2026-06-20 13:36:54 -07:00
  • f8eda1c333 chore(litellm): retire Langfuse — drop success/failure callbacks (redundant + crash-prone) vh 2026-06-20 13:21:31 -07:00
  • 7819f96003 feat(litellm): add gen-frontier / gen-frontier-reasoning aliases (→ GLM 5.2) vh 2026-06-20 10:14:28 -07:00
  • d0eb09cac1 fix(litellm): remove the * → llama-swap wildcard (decommissioned backend) vh 2026-06-20 10:08:20 -07:00
  • d3721034c1 feat(litellm): Worldtree capability aliases (chat-judge, reranker, scalar-judge) vh 2026-06-20 09:53:09 -07:00
  • cd92b85157 feat(omnivoice): tune streaming defaults (16-step + aggressive packing) vh 2026-06-19 22:58:55 -07:00
  • 288d085236 feat(omnivoice): streaming /tts + language-safe sanitizer vh 2026-06-19 22:47:15 -07:00
  • 826c2a6a64 memory: archive 15 pre-2026-06-16 entries to archival-memory.md vh 2026-06-19 21:54:45 -07:00
  • dfda60fac7 memory: snapshot — 2026-06-19 (pt2) litellm task-aliases (classifier->granite, summarizer-large->gen; gen-nt/gen-reasoning-nt added-then-removed as redundant with strip_empty_tools) + gateway-chat model-smoking web chat enhanced (auto-discover /v1/models + image upload) and stood up as a PERSISTENT nginx container on ana-docker :8091 + pi on nh3-dev wired to gen (vision, ~/.pi models.json + gen launcher, local box config) + foot-guns: litellm config-loaded models can't be hot-removed (/model/delete is DB-only; /model/new live-adds work no-bounce but dup on restart) and the * wildcard routes stale/typo'd names to decommissioned llama-swap -> misleading 'Connection error' not 'model not found' (bit a brokkr call to the renamed-away qwen-image-judge). vh 2026-06-19 18:51:03 -07:00
  • 740bcae45d feat(gateway-chat): persistent static-serve stack for the model-smoking web chat vh 2026-06-19 12:32:51 -07:00
  • ef45f6d826 feat(litellm): add classifier -> granite + summarizer-large -> gen aliases (operator) vh 2026-06-19 12:08:26 -07:00
  • 4c40b9fac6 feat(tools): gateway-chat.html — auto-discover gateway models + image upload for vision smoke vh 2026-06-19 11:56:20 -07:00
  • 75bd4c3679 remove gen-nt / gen-reasoning-nt litellm records (operator) vh 2026-06-19 11:56:20 -07:00
  • 2e5ab72e2c feat(litellm): add gen-nt / gen-reasoning-nt (noop-tool + tool_choice:none compat variants) vh 2026-06-19 11:35:48 -07:00
  • 378261763c memory: snapshot — 2026-06-19 gen model = Qwopus3.5-122B vision-intact NVFP4 LIVE on ana-ml2 GPU 0 (full 256K @ fp8 KV + CUDA graphs, util 0.95 + expandable_segments, 92.7 tok/s warm, 3.32x concurrency, text+image+video, tool-calling qwen3_coder; nightly+turboquant-4bit-KV proven UNNECESSARY — stable fp8 reaches 256K) replacing the bjk110 text-only qwen3.5-122b (which displaced mistral-small-4 → Worldtree character backend DARK until repointed, operator-acknowledged) + qwen-image-bench T2I judge replaced qwen3.6-35b-a3b on GPU 1 (alias image-judge) + TP=2 across both Blackwells REJECTED (PCIe-only PIX, no NVLink → all-reduce-bound, one-model-per-card is optimal; PP=2 only if a >96GB model is ever wanted) + foot-guns: MoE FusedMoE workspace is the ~3.1GB un-budgeted floor (can't fill to 0), discard cold tok/s reads (24.8 cold vs 92.7 warm). vh 2026-06-19 11:15:32 -07:00
  • 5b06514020 docs(litellm): gen records now describe Qwopus3.5-122B (vision-intact), not bjk110 text-only vh 2026-06-19 10:25:45 -07:00
  • 20e796cf6b feat(qwopus3.5-122b): gen model → Qwopus3.5-122B vision-intact NVFP4, full 256K @ fp8 vh 2026-06-19 10:24:34 -07:00
  • a5b626b3d5 fix(qwen3.5-122b): enable tool-calling (--enable-auto-tool-choice --tool-call-parser qwen3_xml) vh 2026-06-19 01:56:56 -07:00
  • 5dfce049f4 rename(litellm): qwen-image-judge alias -> image-judge vh 2026-06-19 01:43:25 -07:00
  • bfae924048 feat(qwen-image-bench): replace qwen3.6-35b-a3b on GPU1 with the T2I judge (NVFP4) vh 2026-06-19 01:40:59 -07:00
  • 3ba0e544db tune(qwen3.5-122b): gpu-mem-util 0.90->0.95, max-num-seqs 4->8 (KV 260K->446K tokens, 3.4x concurrency @131K, no OOM) vh 2026-06-19 01:21:28 -07:00
  • 89c83c4271 feat(qwen3.5-122b): replace mistral-small-4 as gen (abliterated NVFP4, text-only) vh 2026-06-19 00:49:04 -07:00
  • 67102b5b94 feat(litellm): add model aliases summarizer / gen / gen-reasoning vh 2026-06-18 23:56:05 -07:00
  • 91688a234b revert(litellm): remove mistral-medium-3.5 entry (GPU0 reverted to small-4 heretic) vh 2026-06-18 23:50:34 -07:00
  • 981ae4e6a1 feat(omnivoice): expose full generation surface (voice-design, language, diffusion params) vh 2026-06-18 23:25:39 -07:00
  • 71f5784016 feat(litellm): add mistral-medium-3.5 (RecViking NVFP4 :8012, temporary GPU0 tenant) vh 2026-06-18 23:12:31 -07:00
  • 06eb487a26 feat(omnivoice): wire to asset-engine via FastAPI wrapper + reuse chatterbox voices vh 2026-06-18 23:03:20 -07:00
  • 984b72757f feat(omnivoice): new TTS stack — k2-fsa/OmniVoice on irv-ml1 3090 vh 2026-06-18 22:25:54 -07:00
  • 715a68bee7 feat(comfyui): native --use-sage-attention (node path dead on 0.24.1) vh 2026-06-18 14:16:56 -07:00
  • 632124c8fb memory: nh3-extdev pi-on-GLM-5.2 wired (mark in-flight item done + residual gates) vh 2026-06-18 14:05:02 -07:00
  • 527a844714 feat(nh3-extdev): install pi (earendil-works) + wire /opt/externs client agents to GLM 5.2 vh 2026-06-18 14:04:29 -07:00
  • a67d4950d0 memory: snapshot — 2026-06-18 heretic abliterated Mistral Small 4 NVFP4 BUILT + LIVE as mistral-small-4 (in-house quant device_map=cpu → native-format convert → drop-in stack under same served-name, A/B'd vs official, operator "heretic stays"; byte-equivalent to official NVFP4) + irv-ml1 VRAM consolidation (ComfyUI pinned to A6000 exclusive/48GB, audio zoo→3090, downed dia/ace-step/csm, comfy-dev torch-pin DISABLE_UPGRADES@2.12.1 + SageAttention rebuilt) + ComfyUI 9-node accel set installed for comfy-dev + ana-ml2 durable vm.overcommit_memory=1 + GLM5.2 wired + nh3-extdev sudo-less manager box + /opt/externs pi-on-GLM client workspaces. vh 2026-06-18 13:48:12 -07:00
  • a8550ad4bc feat(irv-ml1): pin comfyui to A6000 + torch-pin; parakeet -> 3090 (VRAM consolidation) vh 2026-06-18 11:07:43 -07:00
  • f566f61b24 feat(stacks): mistral-small-4-heretic drop-in (abliterated NVFP4 backend swap) vh 2026-06-17 22:33:29 -07:00
  • dd3a5c93fd feat(tools): Mistral Small 4 NVFP4 build pipeline (quant + HF->native converter) vh 2026-06-17 22:19:58 -07:00
  • fc88eff06e feat(ana-ml2): durable vm.overcommit_memory=1 sysctl playbook vh 2026-06-17 22:19:58 -07:00
  • a841eab3ff servers: register nh3-extdev (sudo-less infra-ops manager box) vh 2026-06-17 14:57:06 -07:00
  • fe77a3596a litellm: wire GLM 5.2 (glm-5.2 + glm-5.2-reasoning) via z.ai passthrough vh 2026-06-17 08:56:16 -07:00
  • 8cca365b78 memory: correct gitea action-log API note (per-job endpoint works) vh 2026-06-16 14:56:12 -07:00
  • 03358dccd1 arbo: mount repo pyproject.toml ro into the engine (catalog_version observability) vh 2026-06-16 14:42:04 -07:00
  • 3a7236d51f playbook: put uv/uvx on the irv-ml1-arbo runner PATH vh 2026-06-16 14:24:08 -07:00
  • 47c30e85f9 memory: snapshot — 2026-06-16 litellm strip_empty_tools hook (d1bea13) + single-file gateway-chat.html playground (984ca3d) + claude-bot ADMIN on vh/arbo (arbo CI/CD via service account) + LitBench-RM reward judge served on irv-ml1 A6000 then taken to on-demand (held comfyui's slot) + #295 recall root-cause FLIPPED (score_breakdown-shape DISPROVEN → cold-recall agent_self scope axis vs ratatoskr conjunctive INV-005; Worldtree #297); lessons: litellm-500-Router.acompletion-missing-messages = a request missing Content-Type (NOT a gateway outage — cost 4 needless restarts), litellm-admin-UI-playground-cant-test-vLLM (#6228 empty-tools, proxy-hook-cant-reach-in-process-call), gitea-run-looks-like-never-fired-but-fired-then-skipped/failed-fast (check run list not runner). Archived the 06-09→06-13 cluster (24 entries: 17 decisions + 7 foot-guns). vh 2026-06-16 14:14:36 -07:00
  • 984ca3d383 feat(tools): single-file gateway chat playground vh 2026-06-16 01:40:30 -07:00
  • d1bea13994 fix(litellm): strip empty tools:[] before forwarding to vLLM vh 2026-06-16 00:43:57 -07:00
  • f9277f5440 memory: snapshot — 2026-06-16 ratatoskr Tier-3 MEMORY plane wired (allowlist :8391, key reused, persist+dispatch GREEN; recall-injection root-caused to the score_breakdown shape seam → worldtree-dev #295) + infra-ops durable admin on corviduo (ssh alias + ssh-target) + demo/personal character model qwen→mistral (first-bind-is-default reorder, pin-safe recreate); lessons: bifrost-allowlist-is-per-port, promotion-gate=consumer-agent-memory-block-not-agent_self_enabled, WORLDTREE_IMAGE-pin-from-matrix-sibling vh 2026-06-16 00:13:14 -07:00
  • c99aa49cad feat(corviduo): wire ratatoskr memory plane :8391 into personal Worldtree bifrost allowlist vh 2026-06-15 23:22:06 -07:00
  • aeea377749 memory: snapshot — 2026-06-16 ana-ml2 dual-NVFP4 reshape (GPU0 Mistral Small 4 256K/v0.22.0-vision + GPU1 qwen36 FP8→NVFP4 + Selene FP8 judge + GPU1 grows) + NVFP4-MoE-loads-on-0.23.0 (supersedes blocked) + claude-bot service account (corviduo-org tabled) + arbo→comfy-dev ownership + gitea runner on irv-ml1 + Worldtree demo/personal capability-profile migration (pre-sync-first); lessons: vLLM-0.23-breaks-Mistral-vision (#44911), Mistral-TTFT=Triton-JIT-spikes, vh-user-not-org blocks scoped package-write, old-baseline-instances-need-full-config-set; archived the 2026-06-05/08 cluster (11 entries) vh 2026-06-15 22:26:11 -07:00
  • e124a2f233 tune(gpu1): grow selene 0.13→0.17 + qwen36 0.32→0.34 into the buffer vh 2026-06-15 20:36:04 -07:00
  • c985ede07b feat(selene+mistral): restore Selene judge (FP8, GPU1) + push Mistral to 256K vh 2026-06-15 18:52:26 -07:00
  • 9a49963d07 feat(mistral-small-4): pin v0.22.0 for working VISION baseline + reasoning entry vh 2026-06-15 17:56:37 -07:00
  • c77a9aa4d8 feat(mistral-small-4): deploy NVFP4 119B MoE on GPU 0 (text-only) + gateway vh 2026-06-15 17:28:15 -07:00
  • c6d76051a4 feat(qwen36-vl): swap FP8→NVFP4 + GPU1 rebalance (granite restored) vh 2026-06-15 17:28:15 -07:00
  • 6de0844323 feat(qwen36-vl): split thinking — non-thinking default + qwen3.6-35b-a3b-thinking variant vh 2026-06-15 13:55:11 -07:00
  • 0943d145fb memory: snapshot — 2026-06-15 (cont.) arbo v0.11.22 engine rebuild + catalog v0.11.23 (curated /workflows footer live) + althing-core v0.14.1 box-wide refresh (monitor lock fix) + comfyui VAE-decode SEGFAULT diagnosis (aimdo 0.4.8 cuda-hooks vs torch cu129/cu130 mismatch, NOT OOM); lessons: comfyui-segfault-not-OOM diagnostic, never blanket-kill peer light-monitors vh 2026-06-15 13:35:50 -07:00
  • 12bcd06442 memory: snapshot — 2026-06-15 ratatoskr affect smoke GREEN (Heimdall key mint+inject, allowlist, handshake+emit) + infra-ops bootstrapped on corviduo + dense Qwen3-VL-32B-NVFP4 judge A/B (lost, torn down) + MastMed cloudflared public + R18 clip+caption staged (stub smoke passed; real-voice gate pending) + LiteLLM infra-ops key; lessons: corviduo stale-:latest recreate crash, .claude.json ENOSPC repair, pkill self-match vh 2026-06-15 00:17:44 -07:00
  • 10f346b39e memory: snapshot — 2026-06-14 FP8 vision cutover (qwen35-vl→qwen36-vl, truthful naming, GPU-1 rebalance) + llama-swap pin drop + R16 yield probe executed + standing credential-migration directive; NVFP4-on-vLLM-blocked + sampler-warmup/profiling-race/embed-rerank-waste lessons vh 2026-06-14 14:47:01 -07:00
  • a0fed13801 feat(ana-ml2): replace Qwen3.5-9B vision with Qwen3.6-35B-A3B FP8 on GPU 1 vh 2026-06-14 14:41:49 -07:00
  • b45d0cd86d memory: snapshot — 2026-06-14 arbo auth-off + deploy-pipeline fix (v0.11.6, internal gitea route, catalog-only restart, scripts tracked) + storetank archive decommission (919G -> arbo 502G) + R16 inline arc closed (v1 final); archive 11 (2026-06-04 cluster) vh 2026-06-14 08:56:12 -07:00
  • 6d66bc2f30 feat(arbo): track webhook deploy scripts (arbo-deploy.sh + arbo-webhook.py) vh 2026-06-13 17:17:49 -07:00
  • 6e58e57362 docs(orientation): gitea internal-route gotcha (fleet hosts -> 10.250.50.70:222) vh 2026-06-13 17:04:36 -07:00
  • 5007ec1236 docs(catalog): archive decommissioned + arbo refreshed (502 G post-migration) vh 2026-06-13 15:40:56 -07:00
  • 308ca6f5d2 docs(catalog): record llava_llama3 sweep (919->214 G, 705 G reclaimed) vh 2026-06-13 15:21:11 -07:00
  • 1902425682 docs(catalog): record storetank image-models curation + remaining inventory vh 2026-06-13 15:18:21 -07:00
  • db97899037 feat(arbo): disable ENGINE_TOKEN bearer auth on prod (WireGuard = boundary) vh 2026-06-13 14:05:07 -07:00
  • f32c6ddaab docs(arbo): GRANITE_KEY scope now granite + qwen-vision (extended) vh 2026-06-13 13:45:32 -07:00