3b7e10cd29
Current state rewritten to session-end: secrets-broker LIVE (secret CLI + 25-item backfill + rm/warn), stonehenge-park /park service repo stood up (vor-plan+vor-ui), dots-tts extracted to tts-stack; NEXT PROJECT = wgtunnel (WireGuard over a plane's wifi). Two detail files added; 2 oldest T&A entries archived. Handoff written to /tmp/infra-ops-handoff.md for the post-clear pickup.
1763 lines
139 KiB
Markdown
1763 lines
139 KiB
Markdown
# Archival memory — eshpfi-management
|
||
|
||
_Entries moved out of persistent-memory.md to keep the active file scannable. Read this when researching historical decisions or revisiting past foot-guns._
|
||
|
||
## Recent decisions (archived)
|
||
|
||
- `[2026-05-12]` corviduo-dev (Worldtree-team dev VM, 10.250.50.152,
|
||
CT 106 on pfi-pve) added to `servers/` inventory. Treat like SF
|
||
client hosts: PFI hosts + provides emergency-ops backstop;
|
||
Worldtree team owns OS config + deploys + backup decisions.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-12]` Worldtree `:latest` tag drift bug — fixed by
|
||
health-gated `:latest` advance in vh/worldtree's deploy workflow
|
||
(architect commit `8ef3801`): only tag `:latest` AFTER the new
|
||
container's `/health` probe passes. Build-on-host stacks here
|
||
don't have this problem because the playbook always builds the
|
||
SHA-tagged image from a `git reset --hard <ref>` checkout.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-12]` asset-engine stack scaffolded LAN-direct at
|
||
`http://10.250.50.70:8200`. Initially included Traefik labels for
|
||
public hostname; user pulled them out (internal tool, no public
|
||
TLS surface needed). Pattern: internal tools default LAN-direct;
|
||
Traefik wiring only when external/TLS required.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-12]` asset-engine catalog gains `lifecycle: { stack,
|
||
vram_gb, gpu_device_id }` per irv-ml1 service for the
|
||
orchestrator feature. SSH keypair scaffolded at
|
||
`ana-docker:/opt/docker/conf/asset-engine/ssh/` for
|
||
asset-engine container → irv-ml1 orchestration via dedicated
|
||
ed25519 key.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-13]` `pull-hf-repo.yaml` is the canonical HF-fetch
|
||
playbook on ana-ml2. Supports `--var repo_type=model|dataset|space`.
|
||
Replaces ad-hoc `huggingface_hub.snapshot_download` calls.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-13]` Selene-1-Mini-Llama-3.1-8B added to llama-swap as
|
||
judge model. mradermacher `i1-Q6_K` imatrix quant (~6.5GB). AtlaAI
|
||
reward/eval model — temp 0.01, ctx 32K, q8_0 KV cache. New JUDGE /
|
||
EVAL MODELS section in `stacks/llama-swap/conf/config.yaml`.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-13]` `vllm-qwen3` → `vllm` stack rename. Added
|
||
`vllm-reward` service (Skywork-Reward-V2-Llama-3.1-8B-AWQ classifier).
|
||
Three vLLM services share GPU 1 (embed 0.20, rerank 0.20, reward 0.30
|
||
utilization; 30% headroom). All use `--runner pooling`; classification
|
||
drives via model's `architectures: [LlamaForSequenceClassification]`
|
||
in config.json, NOT `--task classify` (deprecated in vLLM 0.19.1).
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-13]` /tend-docs first pass deletions: `stacks/infinity/`
|
||
removed (retired by vllm). Archived
|
||
`docs/asset-engine/design-brief.md` → `docs/archive/asset-engine/`
|
||
with archival header. Fixed `pfi-pve` VM list to full
|
||
`qm list` enumeration. Dropped stale weak-password section from
|
||
`pfi-postgres` (rotation done 2026-04-23).
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-14]` althing-chamber stack scaffolded: chamber + forseti.
|
||
Internal LAN-only at port 7881 (chamber default 7878 collides with
|
||
task-board). Two-service compose, shared SQLite bind-mount,
|
||
build-on-host pattern via vh/althing's gitea-workflow. Forseti is the
|
||
canonical dev for this stack (galdrabok is on a different project).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-16]` althing-chamber Phase 2: added `althing-agent-runner`
|
||
as third compose service (worldtree-driver agent dispatcher).
|
||
All three althing services use the same image; `command:` selects
|
||
entrypoint. Safe to enable preemptively (sleeps when no
|
||
driver=worldtree handles declared).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` Phase 3.1 cross-process streaming uses Valkey 8
|
||
alpine as a sibling compose service in `stacks/althing-chamber/`,
|
||
redis-protocol pub/sub for high-volume `msg_delta` / `msg_thinking` /
|
||
`msg_start` / `msg_complete` event kinds. DB bridge keeps
|
||
`msg_curated` + `floor_grant` (structured / canonical). Two-channel
|
||
architecture, no overlap. chamber + agent-runner `depends_on:
|
||
valkey: service_healthy`.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` Worldtree admin workflow shift (per vh): infra-ops
|
||
gets its own permanent admin-tier key (`61419c92`, stored at
|
||
`ana-docker:/opt/docker/conf/.secrets/worldtree-infra-ops-admin`).
|
||
Future admin ops route through this key, not the bootstrap admin
|
||
via docker-as-root.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` Worldtree env-var addition checklist: anytime
|
||
introducing `os.environ.get("FOO")` in worldtree code, update BOTH
|
||
`.env.example` AND `compose.yaml`'s `&worldtree-env` anchor in the
|
||
same PR. Same Z_AI_API_KEY-shape footgun bit
|
||
`BIFROST_CLIENT_ALLOWED_HOSTS` (#170) until worldtree-dev added the
|
||
passthrough line in `08f02b2`.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-18]` Volva systemd install complete after three-stage
|
||
debug. Final unit at `/etc/systemd/system/volva.service` runs as
|
||
`User=lkraven` with `ProtectHome=read-only` +
|
||
`ReadWritePaths=/home/lkraven/.althing /home/lkraven/.codex`
|
||
carve-outs for state writes. `VOLVA_ALTHING_CLI=/home/lkraven/
|
||
.local/bin/althing-cli` + `ALTHING_HANDLE=volva` both pinned in
|
||
env.sh.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Worldtree CD disk-hygiene strategy: watermark gate
|
||
(env-tunable threshold + window, fail-loud on still-low post-prune)
|
||
+ eager post-deploy prune (only after `:latest` advance succeeds,
|
||
uses `docker image prune -a --filter "until=24h"` which respects
|
||
in-use semantic — protects pinned + personal images automatically).
|
||
Combined: demo VM holds ~24h of deploy history instead of unbounded
|
||
accumulation. Shipped in vh/Worldtree PR #184 (`306cd61` +
|
||
`613dac2` + `bd91df5`).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Skaldsong CD shape: shape (1) of three operator
|
||
options — container + Gitea registry + pull-restart, matching
|
||
Worldtree's pattern. Target host ana-docker (NOT nh3-dev where
|
||
skaldsong-dev runs for iteration). SHA-pin only for now; health-gated
|
||
`:latest` advance is a follow-up once `/health` exercises Worldtree
|
||
+ Kokoro reachability.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Skaldsong prod (ana-docker) switched from demo
|
||
Worldtree (:8080) to personal (:8081). Same `user_id=skaldsong`
|
||
as the nh3-dev hand-launch key — shared Heimdall agent slot
|
||
(`skaldsong:wizard-v2`), different `key_id`s for independent
|
||
rotation. Demo Worldtree stays for isolation; personal becomes
|
||
the multi-consumer dev iteration instance.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` mead-hall Bifrost v0.3 end-to-end smoke green.
|
||
Closed task #32 (althing thread `01KRV1M2KW6N6HBEXGTH72QXCA`).
|
||
Wire layer (handshake + binding + dispatch) + data-flow
|
||
(per-dispatch JWT claims → `ctx.session_id` populated → real
|
||
session-scoped data) + agent-loop (LLM reads + quotes back) all
|
||
proven. Resolves the "stalled mid-Worldtree" state from the
|
||
2026-05-17 snapshot.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-25]` v0.25.3 lofn tuning: `temperature 0.6 → 1.0` +
|
||
`repetition_penalty 1.0 → 1.15` on default+fast profiles.
|
||
Heretic-abliterated qwen3.6 was locking into degenerate attractors
|
||
at the model's thinking-mode floor (0.6). Pattern:
|
||
abliterated/uncensored Qwen variants need higher temp + non-trivial
|
||
rep-penalty than base, NOT the model-card's documented floors.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-25]` Worldtree #205 v0.25.2 ships `/app/config/` as
|
||
bind-mount + root-then-drop entrypoint shim (gosu). Operators
|
||
get persistent per-instance config without container-rebuild.
|
||
Same bind-mount pattern hit twice subsequently in v0.27.0 (selene)
|
||
and v0.29.9 (echo) — bind-mount shadows image-baked defaults, so
|
||
every new required key surfaces as a crash-loop on existing
|
||
deployments. The v0.29.12 canonical example files close this loop.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-26]` Worldtree v0.27.0/v0.27.1 fixes Tier 3 GET
|
||
visibility. `available_agents()` helper was over-applied to
|
||
`GET /agents/<id>`, masking ALL Tier 3 agents regardless of row
|
||
state. Bug only visible as "agent not found" via GET; storage was
|
||
fine (silent-2xx PATCHes had persisted correctly). v0.27.1 added
|
||
fail-fast hardening for the startup pre-resolve fragility class.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-26]` Skaldsong v0.30.7 defensive 409→PATCH fallback.
|
||
v0.30.6's GET-then-define-or-patch path crash-looped against
|
||
pre-v0.27.0 Worldtree's GET-visibility bug (GET 404 phantom →
|
||
define 409 conflict). v0.30.7 catches the 409 and falls through
|
||
to PATCH (which silently 2xx'd on the pre-fix Worldtree).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-27]` Worldtree v0.29.x landed full saga→echo refactor +
|
||
config-validator hardening (v0.29.10 create_provider family-before-regex;
|
||
v0.29.11 collect-then-raise echo startup validators; v0.29.12 ships
|
||
providers.yaml.example/defaults.yaml.example canonical configs; v0.29.13
|
||
reasoning_content extraction + catalog family lookup). Operator-asked,
|
||
worldtree-dev-shipped, CI-deployed.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` artemis-31b-v1i added to llama-swap + worldtree personal.
|
||
BeaverAI Gemma 4 31B Q6_K (~28.6GB), 128K ctx, `--reasoning-format deepseek`
|
||
(gemma format unsupported in deployed llama.cpp). Worldtree catalog
|
||
`family: gemma` so GemmaProvider routes reasoning tokens.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` Skaldsong streaming TTS v0.32.0→v0.32.2: chunked-batch SSE
|
||
(one Kokoro POST per paragraph); defensive event_stream catch-all; NDJSON
|
||
parsing for Kokoro /dev/captioned_speech multi-line responses.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-31]` Dia2 deployed as **two fixed-model instances**
|
||
(`dia2-2b` :8200, `dia2-1b` :8202) from `local/dia:v2`, retiring
|
||
legacy Dia 1.6B; catalog `dia` entry **removed → dia2-2b + dia2-1b**
|
||
(breaking for asset-engine). Rationale: the devnen wrapper is
|
||
single-model and IGNORES the OpenAI `model` field (verified on its live
|
||
OpenAPI), so the only way to offer both Dia2 models as real per-request
|
||
asset-engine choices is one fixed endpoint per model. `3139e81`
|
||
(deploy), `db15638` (catalog swap).
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` Both dia2 catalog entries route to the wrapper's richer
|
||
**`/tts`** endpoint (not `/v1/audio/speech`) to expose the full control
|
||
surface (cfg_scale/temperature/top_p/cfg_filter_top_k/voice_mode/clone);
|
||
all defaults sourced from the wrapper's `CustomTTSRequest` Pydantic
|
||
blessed values. Voice default is `voice_mode: clone` +
|
||
`clone_reference_filename: Abigail.wav` so a stable (non-random-gender)
|
||
voice is the out-of-box behavior. `55602b7`, `5c47843`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` Zonos REST adapter (`stacks/zonos/adapter/`,
|
||
`local/zonos-api`) — thin OpenAI-ish `/v1/audio/speech` FastAPI in front
|
||
of the Gradio-only Zonos SDK; JSON-envelope `{audio, audio_format,
|
||
seed}` (Zonos is the fleet's first seedable TTS). Port 8203 (moved off
|
||
8201 — collided with csm). Built; NOT deployed (stack down for VRAM).
|
||
Also fixed the upstream image's missing CMD (`71df6f7`). `81efa8d`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` Catalog schema regenerated: added `CatalogLifecycle` +
|
||
`reproducibility.seed_field` (`b7b2130`). Resolves the stale-schema
|
||
hand-off; catalog now validates clean. (asset_engine `catalog.py`
|
||
Pydantic already supported both — schema file was just behind.)
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` TTS bench expanded with `stacks/{dia,zonos,csm}`
|
||
(`666f7f3` dia+zonos, `a4b8c2a` csm). The bench already had Fish
|
||
S2-Pro / Chatterbox-Turbo / IndexTTS-2 / CosyVoice3 / Kokoro /
|
||
VibeVoice / Qwen3-TTS / Kyutai. (csm since removed 2026-06-01.)
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` Remote browser/iPad/Vision-Pro driver seat for the
|
||
agent-fleet zellij session `Claude` stood up on nh3-dev (ttyd behind
|
||
Caddy, network-gated). Out of this repo — full architecture + the
|
||
HTTP2/OSC52/Safari-auth gotchas in auto-memory `reference_ttyd_fleet_seat`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-30]` esh-docker-vm NFS boot-ordering fix:
|
||
`playbooks/fix-esh-nfs-boot-ordering.yaml` (`c0458d9`, + `53157b1`
|
||
drop-in filename-collision fix) adds `_netdev,nofail` to the four
|
||
10.0.50.50 NFS mounts + a docker `After=remote-fs.target` drop-in —
|
||
resolves paperless `Exited(255)` on reboot. traefik also gained
|
||
`restart: unless-stopped`. Full incident → auto-memory
|
||
`incident_esh_docker_nfs_boot_race`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-06-01]` **Fish cloning VERIFIED competent (ECAPA-TDNN)** — retracting
|
||
the earlier "weak cloner" call. Isolated test: Imogen-referenced clone ~0.79
|
||
cosine to the real `Imogen.wav` vs ~0.10 for the no-reference default;
|
||
transcript condition (correct 0.787 / bogus 0.778 / empty 0.738) barely moves
|
||
identity (affects pronunciation, not timbre). Root cause of "Imogen sounds
|
||
nothing like British" = the `"undefined"` select bug feeding Fish its default
|
||
voice, NOT Fish. So the entire Fish-Imogen saga was the undefined bug; no
|
||
Fish-side fix needed. (Methodology lessons → Tried and abandoned.)
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **CSM (Sesame csm-1b) torn down entirely** — removed from
|
||
catalog, `stacks/csm/`, `playbooks/deploy-csm.yaml`, and host
|
||
(`c54ab13`). Two reasons: (1) deep-research verdict — the acclaimed
|
||
Maya/Miles demo is a fine-tuned, larger variant Sesame never open-sourced;
|
||
the open 1B base ships no usable voices, can't generate text, English-only,
|
||
no real-time streaming, absent from leaderboards, dominated by
|
||
Kokoro/Dia2/Fish/IndexTTS. (2) The upstream image is rotted (see Tried and
|
||
abandoned). asset-engine re-synced (`306aae9`).
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Chatterbox catalog v1→v2** (`93cedd6`) — re-point to the
|
||
devnen wrapper's `/tts` (was the thin OpenAI `/v1/audio/speech` exposing
|
||
none of the emotion knobs → poor output). Default `exaggeration: 0.5` per
|
||
Resemble's README ("works well for most"); the server SHIPS demo-tuned 1.3
|
||
which over-acts — deliberately NOT adopted. 28 predefined voices via
|
||
`/get_predefined_voices` (default Emily.wav). Corrected seedable false→true.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Fish-S2 catalog v1→v6** (`f3c260e`..`38d9e3b`) — was a
|
||
blank 2-field form; now exposes the full fish-speech ServeTTSRequest surface
|
||
+ a 32-voice `reference_id` "Voice" dropdown (STATIC list — Fish has no
|
||
/voices API). `reference_id`-by-name works ONLY when each voice is staged as
|
||
a `references/<name>/` dir with `<name>.wav` + `<name>.lab` — was silently a
|
||
no-op until that layout fix on 2026-06-01 (see the two-root-causes entry
|
||
above); the inline-path-as-filesystem-path `references` shape 500s (see Tried
|
||
and abandoned). Corrected vram
|
||
4→24 GB, cold_start 8→240s, seedable false→true. Added 3 consenting
|
||
British-female VCTK voices (Imogen=p225/Eleanor=p228/Beatrice=p229).
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **`blendable: false` catalog flag — DEFERRED**, pending
|
||
asset-engine-dev honoring it in the form template. Clone-voice selects
|
||
(chatterbox/fish/dia2) must render as plain dropdowns, not the Kokoro blend
|
||
widget (blending reference clones is nonsensical). I ship the catalog flag +
|
||
schema field once they read it. Tracked: althing msg
|
||
`01KT2K2SY9N7AY69R9V0B4RXSW` to asset-engine-dev.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Fish reference_id empty-dir fix shipped** (`c5bbb90`) — Populated
|
||
`references/<name>/<name>.wav`+`.lab` for all 32 voices; playbook gained
|
||
normalize-step + A/B smoke gate. glados got a real transcript (ASR'd via
|
||
Parakeet): the Portal "Welcome to test chamber 4" lines.
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-02]` **Chatterbox → main TTS engine; build custom `chatterbox-fast`
|
||
streaming container.** Workload = single-stream interactive. GPU placement:
|
||
3090 (device 0) if it fits else A6000 (device 1). Cutover: parallel catalog
|
||
entry, burn in beside live `chatterbox`, then flip. Streaming = adaptive
|
||
buffer-ratchet chunking (native frame-streaming abandoned). Since extracted to
|
||
its own repo `vh/chatterbox-fast` (auto-memory `reference_chatterbox_fast_repo`).
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-02]` **Sentence-splitting loses quality (operator-corrected).** Claimed
|
||
naive sentence-level streaming has "zero quality loss" — WRONG. The T3 AR backbone
|
||
conditions prosody on the WHOLE text; splitting loses cross-sentence prosodic
|
||
context even though voice timbre stays (reference-conditioned). No *artifacts* ≠
|
||
no *quality loss*. Hence adaptive-chunk design (maximize context per chunk subject
|
||
to latency budget), not fixed per-sentence splits.
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-03]` **Declined worldtree v0.30.4 staging deploy** — that's worldtree-team's
|
||
CI/CD lane (a developer `staging/vX.Y.Z` git-tag promote), not infra-ops. They self-corrected
|
||
to the same conclusion independently.
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **R14 scope = (a) provision-only.** infra-ops provides box + CUDA env +
|
||
engines + weights + NFS; brokkr/dev wires `arms.py` + runs — keeps infra-ops OFF the
|
||
VIVAE-processing path (VIVAE = Variably Intense Vocalizations of Affect/Emotion, CHARTER §4
|
||
highest-liability; operator authorized R&D-eval-only, quarantined). Box = irv-ml1 (A6000
|
||
free; ana-ml2 GPU-saturated). Per-engine venvs (divergent torch stacks); A6000 = `cuda:0`
|
||
NATIVE (≠ docker `=1`).
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **yt-voice-clipper push-to-deploy via gitea webhook** (operator-directed,
|
||
after 6 manual rebuilds in ~40 min). Webhook (not poll) — gitea CAN reach the WG IP per the
|
||
operator. The proxy env + Homepage labels live in the **host-specific override** (untracked
|
||
→ survive the auto-deploy's `git reset --hard`), NOT yt-voice-clipper-dev's image. Runbook
|
||
`d4f180d`.
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **yt-voice-clipper bot-gate fix = route yt-dlp through NH3 residential
|
||
egress, NOT cookies/PO-token.** YouTube hard-flags the Irvine colo IP (LOGIN_REQUIRED on a
|
||
public video even with no cookies). Cookies + the bgutil PO-token + deno JS-runtime all
|
||
loaded fine — the gate is pure IP reputation. Operator chose proxy-via-nh3-dev → durable
|
||
dante proxy → proven. The egress proxy is a reusable fleet lever for any datacenter-IP-gated
|
||
service.
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-04]` **`brokkr` user (no-sudo) on irv-ml1; R14/R15/R16 substrate moved to /home/brokkr.** Persistent box services there need SYSTEM systemd units (see Tried).
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **ollama upgraded 0.9.0→0.30.4 on irv-ml1** (Ministral-3 is a Dec-2025 model the old engine refused); A6000 pinned by **UUID** not index (native fastest-first ≠ nvidia-smi PCI).
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **Worldtree demo/pinned/personal deploys are ALL CI/CD, not infra-ops** — a "deploy vX.Y.Z" request to infra-ops is MISROUTED → point them back to their pipeline. (`d8d776c`, auto-memory `reference_worldtree_deploys_cicd`)
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **infra-ops NOPASSWD-sudo identity commissioned, scoped to PFI boxes** (+esh-docker-vm by operator override) — so infra-ops completes DevOps end-to-end vs handing the operator sudo steps. Dedicated key, sudo log_output, key-gated. (`8c32a05`, auto-memory `reference_infra_ops_sudo_identity`)
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **phi4 ships the CANONICAL/official Phi-4 chat template, NOT Ollama's** (Ollama's omits the system `<|end|>`). Applied an Ollama-matching override then reverted — ship correct, not the benchmark quirk. (`90e08f0`→`27eb537`; "headgun" lesson in Tried.)
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **phi4-mini FP8 on ana-ml2 vLLM is the nevermore summarizer; granite-4-small retired** from llama-swap. 50K ctx + FP8 KV. (`40a374b`) [superseded by the 2026-06-05 granite-4.1-8b decision]
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-08]` **volva.service + heid.service removed from nh3-dev** — vestigial systemd daemons; Heid/Volva re-architected from Python pollers to Claude Code session orchestrators (heid `12aa5a9`); volva.service was crash-looping 203/EXEC. (`6e2f80e`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Granite 4.1 8B FP8 replaced phi4-mini as the production summarizer.** Beat phi4 on precision in brokkr's R15 P03. **Staying FP8, not Q4/AWQ** — primary workload is high-concurrency, where FP8 scales ~linearly (2010 tok/s @ C=32). vLLM `vllm-granite` :8004 GPU 1, official IBM compressed-tensors FP8. (Then on Ada; box has since gone Blackwell.) (`34a43a0`, auto-memory `reference_ana_ml2_vllm_granite`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Langfuse v3 on ana-docker (:3001) as the gateway trace UI**; LiteLLM `success_callback:[langfuse]` live. Pretty traces + tok/s dashboard. NOT a prerequisite (spend_logs already capture tokens+latency). (`9171e6a`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Ollama BANNED fleet-wide** (operator directive) — never stand one up; tear down any found; serve via llama-swap or vLLM. (auto-memory `feedback_avoid_ollama`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **ComfyUI / FLUX.2 work split to `~/development/comfy-dev`** (dedicated repo + agent). eshpfi keeps the `comfyui`/`arbo` stack compose; comfy-dev owns the model/workflow knowledge. (auto-memory `reference_irv_ml1_ampere_quant`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Worldtree summarizer config refresh DEFERRED to Worldtree #254** (granite-4.1-8b is the structured-output profile, ON HOLD, no live consumer). Bind-mount CONFIGS (providers.yaml, vh-owned) ARE infra-ops's to apply directly — only the app/image DEPLOY is CICD; the `.env` is deploy-owned. (auto-memory `reference_worldtree_deploys_cicd`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Arbo prod bearer auth turned OFF — WireGuard is the access boundary** (operator decision; reverses ADR-0001's "closed the open-auth hole"). ENGINE_TOKEN must be **ABSENT, not empty** (empty-string still gates) — removed from BOTH the host `.env` AND the compose `environment:` injection line. Original token backed up at `irv-ml1:/opt/docker/compose/arbo/.env.pre-auth-off.bak`; comfy-dev updated their ADR-0001. Tracking: `db97899` + `playbooks/arbo-disable-engine-token.yaml`.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Storetank image-models archive DECOMMISSIONED; arbo is the single live ComfyUI model tree (502 G).** Curated `/storetank/image-models/comfy` (was 919 G, the native `/opt/ComfyUI/models` symlink target) → killed everything superseded by arbo's current gen (Hunyuan, WAN2.1, FLUX.1, Chroma, SD3.5, orphaned umt5+llava ≈ 739 G) + migrated the keepers (gen-agnostic utilities + the SDXL/Pony stack, 177 G) into `/storetank/arbo/models` (same-fs move, skip-existing protects prod). Tracking: `docs/storetank-image-models-archive.md` + `docs/arbo-comfyui-model-catalog.md` (`1902425`→`5007ec1`).
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **GRANITE_KEY provisioned to comfy-dev's nh3-dev dev env** at `~/.arbo_granite_key` (0600) for the hero gen+judge script — verbatim copy of the prod `arbo-prompt-enhance` vkey (now extended to reach BOTH `granite-4.1-8b` AND `qwen3.5-9b-fp8`); nothing minted. The vkey README's "granite-only" wording was stale → corrected (`f32c6dd`).
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **ana-ml2 upgraded Ada → dual RTX PRO 6000 Blackwell Max-Q** (96 GB each, cc 12.0 / sm_120; was dual RTX 6000 Ada 48 GB / cc 8.9 — confirmed live via `nvidia-smi`). Unlocks NVFP4 (FP4 tensor cores) and doubles VRAM headroom. CLAUDE.md GPU-spec doc-fix **LANDED `355a240`** (operator). Tracking: `19a07b9`/`1e2a3a1` ("Blackwell 96GB").
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **NVFP4-W4A4 is infeasible for Granite — FP8 stays the Granite-on-Blackwell format.** W4A4 collapses at 30k context, proven **producer-independent** (modelopt AND llm-compressor both clean-NONE from the same BF16 base + wikitext-2k calib). No 4-bit wins both axes: W4A4 = quality collapse; W4A16-NVFP4/AWQ = weight-only dequant → bf16 (no FP4-core speedup). **30B retired**. (auto-memory `reference_nvfp4_w4a4_granite_infeasible`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Qwen3.5-9B VL (FP8) deployed on ana-ml2 GPU 1** — `qwen35-vl` stack, :8007, gateway alias `qwen3.5-9b-fp8`. **Pinned nightly digest, not `:latest`**: the stable release quantizes the VL *vision tower* under `--quantization fp8` → garbage vision (LM fine); the nightly correctly excludes it. Re-pin + drop the pin once that exclusion lands stable. (`2e3dcc2`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **comfyui 325 G model tree migrated worktank → `/storetank/arbo`** (worktank 97% → 26%). `arbo` is the consuming app; overlay bind-mount via `COMFYUI_MODELS_DIR`. (`38186be`) (See the 2026-06-13 archive-decommission decision above — this tree later absorbed the storetank-archive keepers, reaching 502 G.)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **GPU layout settled on the Blackwell box.** GPU 0 held free for large-model hot-loads (llama-swap pinned, `edf0f91`); GPU 1 steady-tenant — granite 131k ctx, qwen 65k, embed/rerank/reward trio, ~3.5 GB free (`1e2a3a1`, `19a07b9`; trio re-floored for 96 GB, 20×-parallel-stable). embed/rerank left at floor — long docs chunked BEFORE embedding. max-model-len is a ceiling not a reservation.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Prefix caching pinned explicit on granite + qwen** — benched ~6.5× faster TTFT on a shared ~4.5k-token summarizer template; soft/evictable, neutral when prefixes don't repeat. vLLM v1 defaults it ON (granite) but the qwen nightly defaults OFF — pin both. (`a9a2be7`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **granite-4.1-8b listed as the always-available summarizer/classifier + a shared all-agents key minted** (operator-directed). Global `~/.claude/CLAUDE.md` Global-tools entry; key alias `all-agents-local`, scoped to the FREE local models only (granite + qwen-vision + embed/rerank, NOT paid GLM), internal-gateway-only, rotatable. (auto-memory `reference_litellm_gateway`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **arbo engine + frontend stack stood up** (ADR-0001) — irv-ml1 co-located inference engine (`ee57e69`), python-based healthcheck (slim image, no curl/wget, `bdb3312`), frontend ro-mounted from the checkout (`922e8ad`, ADR-0001 D2).
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-11]` **GLM thinking inverted at the LiteLLM gateway** (operator call): `glm-5.1` defaults thinking-OFF; `glm-5.1-reasoning` = same z.ai upstream, thinking ON. Mechanism: `litellm_params.extra_body:{thinking:{type:disabled}}` — `drop_params` strips a top-level `thinking`/`reasoning_effort` but forwards `extra_body` verbatim to z.ai. Shared-gateway change. (`95b2701`, auto-memory `reference_litellm_gateway`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-11]` **pi coding agent installed on nh3-dev as a GLM 5.1 harness** — `@earendil-works/pi-coding-agent` via **bun** (npm's global prefix is `/usr` → needs sudo, bun avoids it). Config `~/.pi/agent/models.json`, launcher `~/.local/bin/glm`.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-11]` **z.ai web-tools (regin) = z.ai hosted MCP path, NOT the `/paas/v4` Tool API.** WORKS: MCP streamable-HTTP at `https://api.z.ai/api/mcp/{web_search_prime,web_reader}/mcp`, `Authorization: Bearer $ZAI_API_KEY` (the **MCP** key, distinct from `Z_AI_API_KEY` the LLM key). Reference impl = Worldtree's Leif agent.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-10]` **Mac Pro migration framed: workstation-only** (M2 Ultra ARM, racked NH3 on-subnet); sidecars stay Linux. `migration-plan.md`. (See in-flight.)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-10]` **Worldtree deployed-config propagation is infra-ops's OWNED lane** (operator ruling). worldtree-dev pings the config delta pre-merge; infra-ops syncs `config/*.yaml` from MERGED canonical to the `/opt/worldtree*/config` bind-mounts on demo+personal. The v0.33.8 9-HOUR demo outage (a `model_roles.yaml` startup-dep that never reached the VM) is the failure mode this prevents. providers.yaml stays hand-tuned. corviduo emergency-ops = `ssh vh@10.250.50.152`, docker no-sudo. (auto-memory `reference_worldtree_deploys_cicd`, `reference_corviduo_dev_emergency_ops`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-09]` **LiteLLM scoped virtual keys issued to consumers** (operator-authorized): `brokkr-smithy` (all-proxy-models), `arbo-prompt-enhance` (comfy-dev — granite, later extended to qwen-vision). Mint via `/key/generate` (master `sk-corvid`), scope-restricted + rotatable, value → 600 file never the bus. (auto-memory `reference_litellm_gateway`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-14]` **ana-ml2 GPU-1 vision upgraded: Qwen3.5-9B → Qwen3.6-35B-A3B (official FP8), served under its TRUE name only.** `qwen36-vl` replaces `qwen35-vl` on :8007 (`a0fed13`). The stale `qwen3.5-9b-fp8` name is KILLED at vLLM AND the litellm gateway (404/400) — a model is NEVER aliased under a prior model's name (silent substitution = downstream footgun; operator directive). Consumer comfy-dev/arbo migrated; arbo vkeys → all-proxy-models; shared `all-agents-local` key repointed qwen3.5-9b-fp8 → qwen3.6-35b-a3b. GPU-1 rebalanced for the ~34 GB FP8 weights (granite 0.35→0.24/64K; embed/rerank 0.05→0.03, reclaimed ~4 GB util-waste). Validated: vision correct, 20-concurrent = no OOM. (auto-memory `feedback_no_false_model_aliases`)
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **NVFP4 was the lighter fit (~21 GB) but is BLOCKED on vLLM — FP8 is the working vision path.** `nvidia/Qwen3.6-35B-A3B-NVFP4` won't load: the ModelOpt-NVFP4-MoE loader errors on expert/lm_head scale keys across 0.19.1 (`w2_input_scale`) AND 0.22.0 (`lm_head.input_scale`, vllm #44081) — a pattern across modelopt NVFP4 MoEs. Revisit NVFP4 (frees ~13 GB on GPU 1) once fixed; the 21 GB checkpoint stays cached on ana-ml2. **(SUPERSEDED 2026-06-16 — it loads on vLLM 0.23.0; qwen36 swapped to NVFP4. See the top of this section.)**
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **llama-swap qwen3.5-9b GPU-0 pin DROPPED; GPU 0 reserved for a creative-writing model (pick DEFERRED by operator).** Deep-research (this session) on big-fast-uncensored creative for a 96 GB Blackwell: **GLM-Steam-106B-A12B** (already in the llama-swap config — balanced default) vs **TheDrummer/Behemoth-X-123B-v2** (prose-tier, tops UGI writing+willingness) vs XORTRON-123B (max willingness, weak prose); GGUF-on-llama-swap is the serving path. Tracking: this session + llama-swap config (GLM-Steam present, `untracked by operator choice`).
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **R16 splice-pivot yield probe executed** (infra-ops ran the irv-ml1 inference for brokkr; brokkr owns design + analysis). See Current state. Tracking: althing thread `01KV010WGS…`, `gen_yield_probe.py` in `irv-ml1:~/r16-vmoan-harness`.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **R16 vmoan inline-generation arc CLOSED — v1 at default decode (rep1.2/temp0.8) is the final Chatterbox-tag inline artifact.** Operator's ear rejected every alternative: v2/v3 windowing (omission vs coherence-loss), v4 multi-tag (cohesion held but lost to capacity-competition), emergent inline-token modulation (degenerates, not modulates), and the gen-time decode-polish sweep (soft tamers cut the NVV itself — same omission family as v2; p0 baseline beat p1). All adapters v1–v4 + `tokenizer.json.v3bak` preserved on `irv-ml1:~/r16-vmoan-harness`. Likely-next direction (deferred, NOT formalized): generate→bin→splice + one-shot-clone NVV pipeline routing around the inline-coherence wall. Tracking: brokkr R16 journal + althing thread `01KV010WGSSMPWRNCPAGSPK15Y`.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **Arbo deploy pipeline fixed, hardened, and version-controlled.** Prod rebuilt v0.11.1 → **v0.11.6** backend; the webhook machinery (`arbo-deploy.sh` + `arbo-webhook.py`, :9009 HMAC listener) is now repo-tracked at `stacks/arbo/` (was host-only = recoverability foot-gun). Deploy reaches gitea via the INTERNAL route (`10.250.50.70:222`) and restarts the engine ONLY on `catalog/` changes (graphs/frontend per-request; warn on `src/`|`Dockerfile` only — pyproject/uv.lock churn every commit). Operator kept arbo stack ownership in **eshpfi** (not migrated to comfy-dev's repo). Secret + `.env` stay host-only. Tracking: `6d66bc2`, `6e58e57`, `stacks/arbo/README` Q5.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-15]` **qwen36 thinking-split at the gateway** (operator) — base
|
||
`qwen3.6-35b-a3b` = thinking-OFF default (the vision judge needs it), `…-a3b-thinking`
|
||
= opt-in reasoning. One hybrid checkpoint + `--reasoning-parser qwen3`; mirrors the
|
||
glm-5.1 split. The judge's enable_thinking=false is what makes the NVFP4 swap
|
||
transparent. (`6de0844`)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **arbo engine rebuilt v0.11.6 → v0.11.22 from main; catalog re-applied (v0.11.23) — curated `/workflows` footer live.** src/ schema changes (Model/Lora.summary, Workflow.blurb/keywords + the /workflows endpoint) need a MANUAL image rebuild (the webhook warns-only on `src/`/`Dockerfile`). Built from `/worktank/arbo/repo`, bumped `ARBO_IMAGE` in the host .env, recreated + verified healthy (auto-rollback armed, unneeded). comfy-dev's v0.11.20 catalog push had first crash-looped prod (old baked schema `extra=forbid` vs new git-mounted catalog fields → ValidationError); they reverted the catalog, so rebuild-from-main loaded clean, then they re-applied (v0.11.23). comfyui is a SEPARATE compose project — the rebuild didn't touch it.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **althing-core refreshed to v0.14.1 box-wide on nh3-dev** (forseti + heid both routed it; central-sequenced per DevOps-delegation since `uv tool install --force` swaps the shared binary all handles use). Fixes the monitor non-truncating-lock bug so `stop-monitor`/`--stop` reliably recovers orphans. The editable install (`uv tool install --force --editable .`) covers every nh3-dev handle; `--version` trued up 0.14.0→0.14.1; verified `--stop` works against a real orphan.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **ratatoskr Tier-3 affect live smoke GREEN** (consumer-confirmed). Brokered/minted consumer_id="ratatoskr" Heimdall HS256 key on personal Worldtree (the key IS BOTH the create-session Bearer AND ratatoskr's JWT verifier secret — plain utf-8 string, NOT base64/raw); injected ENV-ONLY to `~/.config/ratatoskr/provider.env` (mode 600) on nh3-dev; added `10.100.10.50:8390` to personal's `BIFROST_CLIENT_ALLOWED_HOSTS`; drove define→session (handshake 200)→turn→affect.emit (200, snapshot persisted). Worldtree signs Bifrost with the caller's raw Bearer verbatim. Drive commands + claims contract in althing thread `01KV4ZYSE3…`.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **Dense Qwen3-VL-32B-NVFP4 A/B'd as arbo's anatomy hero-judge vs the qwen3.6-35b-a3b MoE → dense LOST, pilot torn down, GPU 0 reclaimed.** Dense NVFP4 loads clean on vLLM where the MoE NVFP4 doesn't (no expert-scale keys); ~44 tok/s on Blackwell GPU 0; vision tower stays high-precision. But comfy-dev's ground-truth A/B: MoE wins (faster + dual-use as general LLM + slight fine-detail edge), and the subtle-proportion tier is below ALL VLs. **Don't re-propose a dense-VL judge swap.** (auto-memory `project-qwen3vl-dense-pilot`)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **MastMed prototype containerized + exposed publicly via cloudflared.** Self-contained static SPA (bun→caddy multi-stage build), basic-auth gated (operator's gating call), ana-docker:8089, cloudflared quick tunnel. See in-flight for the ephemeral-URL caveat + named-tunnel upgrade path.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **infra-ops bootstrapped onto corviduo-dev** (operator-directed) — infra-ops now has SSH + NOPASSWD sudo there, OVERRIDING the fleet bootstrap script's deliberate partner-box exclusion (explicit host arg). corviduo config edits no longer need `vh` + a sudo password. (auto-memory `reference_corviduo_dev_emergency_ops`)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **R18 clip+caption placement = ana-ml2** (operator: "ana-docker… ana-ml2 unless irv-ml1 has room"; irv-ml1 lacked free VRAM for the 30B-A3B captioner). Models staged + stub deployed; see in-flight + the hard real-voice gate.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-14]` **LiteLLM infra-ops admin key provisioned** (operator) — resolves the LiteLLM half of the credential-migration directive; use it for ALL gateway ops (NOT `sk-corvid`). Value at `~/.config/litellm/infra-ops-key` (mode 600); gateway reachable directly from nh3-dev at 10.250.50.70:4000. (auto-memory `reference_litellm_infra_ops_key`)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-16]` **litellm `strip_empty_tools` pre-call hook shipped** (`d1bea13`) — an empty `tools:[]` 500s vLLM ("tools must not be an empty array"); a global `litellm_settings.callbacks` CustomLogger pops it (+ orphaned `tool_choice`) before forwarding, so it covers EVERY vLLM model, not one. `drop_params` only drops unsupported PARAMS, not empty VALUES. Mounts beside config.yaml (litellm resolves callbacks relative to the config dir). Verified live across granite/mistral/stream. (`stacks/litellm/conf/strip_empty_tools.py`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **single-file `gateway-chat.html` playground shipped** (`984ca3d`, `tools/`) — zero-dep browser chat straight to the gateway (`:4000`, CORS open), system-prompt box, streaming SSE, renders `reasoning_content`, NEVER sends `tools`. Built because the LiteLLM admin-UI playground can't test vLLM-backed models (see Tried-and-abandoned). Serve on-request via `python3 -m http.server -d tools`.
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **claude-bot granted ADMIN on vh/arbo** (operator) — arbo CI/CD gitea ops now route through the service account, not the operator's personal `vh` (the credential-migration directive). Operator added claude-bot as Administrator collaborator; claude-bot self-serves via its `gitea-password` (basic-auth = FULL perms; its stored token is package-write-only, can't change repo settings). (auto-memory `reference_claude_bot_gitea_creds`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **ratatoskr Tier-3 MEMORY plane wired + live-smoke driven.** The missing server-side step was the per-PORT `BIFROST_CLIENT_ALLOWED_HOSTS` allowlist (`:8391`; consumer key reused, plane-agnostic — my earlier "reachability is the whole wiring" was incomplete). Persist + dispatch proven end-to-end against personal v0.35.3; cross-session recall-INJECTION root-caused to the `score_breakdown` shape seam and handed to worldtree-dev (#295 recall leg open). Driving a real memory-call needs a consumer agent defined WITH a `memory` block (`ctx.memory_config≠None`), not `agent_self_enabled`. (tracking: c99aa49 + auto-memory `reference_bifrost_plane_wiring`) — **[2026-06-16] FLIPPED:** the `score_breakdown`-shape root-cause was disproven; real cause = cold-recall `scope_filter` adds an `agent_self` axis → ratatoskr's conjunctive INV-005 match drops the `{end_user}`-only chunks → 0 hits. Now Worldtree #297. (auto-memory `project_worldtree_295_recall_rootcause`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **infra-ops = the DURABLE admin path on corviduo-dev** (operator granted standing admin; extends the 2026-06-15 bootstrap) — added a `Host corviduo-dev → User infra-ops` `~/.ssh/config` alias + repointed `servers/corviduo-dev/ssh-target` to `infra-ops@`; `ssh corviduo-dev` / `elway corviduo-dev` now land as infra-ops w/ NOPASSWD (vh = explicit-only human path). elway resolves via ssh-config, NOT the ssh-target file. (tracking: c99aa49 + auto-memory `reference_corviduo_dev_emergency_ops`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **demo + personal `character` model flipped qwen3.6-35-a3b → mistral-small-4** (operator) — in `model_roles.yaml` the FIRST `catalog_id` in a role's `binds:` is the default; reordered mistral-first (qwen retained in the switch-allowlist, non-destructive), applied via the PIN-SAFE api recreate (derive `WORLDTREE_IMAGE` from the matrix sibling — never a bare `compose up`, which falls to the crash-blocking stale `:latest`), verified a fresh character agent resolves model=mistral-small-4. (auto-memory `reference_corviduo_dev_emergency_ops`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **NVFP4 ModelOpt MoE now LOADS on vLLM 0.23.0 — SUPERSEDES the
|
||
2026-06-14 "blocked" finding (#44081 fixed).** Confirmed the cached
|
||
`nvidia/Qwen3.6-35B-A3B-NVFP4` loads clean + generates on 0.23.0; swapped prod
|
||
qwen36 FP8→NVFP4 (~13 GB reclaimed on GPU 1, fp16 KV). Validated by brokkr (speed:
|
||
faster single-stream, ≥parity batched) + comfy-dev (vision anatomy-judge A/B: PASS).
|
||
One W4A4 tell: minor multi-step chained-reasoning slip — doesn't bite the vision
|
||
judge. Mistral's NVFP4 is a SEPARATE compressed-tensors path (day-0), distinct from
|
||
the ModelOpt one. (auto-memory `reference_nvfp4_moe_loads_on_vllm_023`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **Mistral Small 4 = the GPU-0 tenant** (operator) — fills the slot
|
||
the deferred creative-writing pick held, and supersedes llama-swap (now down).
|
||
NVFP4 119B MoE, pinned **vLLM v0.22.0** for working vision (0.23.0 breaks it),
|
||
256K context. Operator's creative-writing model; **abliteration planned** → then it
|
||
becomes the uncensored-creative model the persona substrate needs. Reasoning via
|
||
`reasoning_effort` (BINARY none|high — no medium). (tracking: `c77a9aa`/`9a49963`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **Selene judge restored on vLLM FP8, NOT NVFP4** — W4A4 is too
|
||
aggressive for a precision judge validated at Q6_K; FP8 ≥ Q6_K fidelity, and Selene
|
||
is text-only Llama-3.1 so dynamic `--quantization fp8` has no vision tower to noise.
|
||
`selene-1-mini-8b` on GPU 1 :8011. Worldtree's Domari role repointed to it. (`c985ede`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **GPU-1 buffer put to work: selene 0.13→0.17 + qwen36 0.32→0.34**
|
||
— the two KV-bound generative models that gained Worldtree consumers; granite + the
|
||
pooling models under-use their util, so growing THEM is wasted. (`e124a2f`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **claude-bot Gitea service account stood up; corviduo-ORG conversion
|
||
TABLED** (operator). claude-bot is the credential-migration foundation. The proper
|
||
fix for scoped agent package-write is converting vh (a USER) → a `corviduo` ORG —
|
||
blast radius assessed (org creation = free; arbo-first migration = cheap, no registry
|
||
images; full fleet = incremental, ~8 active vh images / 50 repos) — but tabled for
|
||
later. Registry-push deferred until then. (auto-memory `project_migrate_infra_access_to_claude_credentials`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **arbo engine builds → comfy-dev OWNERSHIP** (operator approved) +
|
||
Gitea Actions runner stood up on irv-ml1 (host-executor as lkraven). infra-ops keeps
|
||
box-level escalation (runner, package:write PAT pending corviduo-org, webhook). See
|
||
the runner entry in Current state. (tracking: this snapshot + arbo thread `01KV70AB2B4B…`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **Worldtree persona substrate migrated qwen→Mistral via capability-
|
||
named profiles** (v0.35.8→v0.35.10; demo + personal). PRE-SYNC-FIRST is now the
|
||
standing pattern for any providers.yaml/profile change: sync config to the bind-mount
|
||
BEFORE the deploy image-swap restart, so new agents + new config land together (the
|
||
v0.35.9 merge-first attempt race'd → 14 crash-loops). (tracking: worldtree thread `01KV6X0M…`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-19]` **litellm task-aliases: `classifier`→granite-4.1-8b, `summarizer-large`→gen/qwopus** (`summarizer`→granite already existed). Also added then REMOVED `gen-nt`/`gen-reasoning-nt` (redundant with the global `strip_empty_tools` hook). (ef45f6d, 2e5ab72, 75bd4c3)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **gateway-chat: model-smoking web chat → PERSISTENT container.** `/v1/models` auto-discovery + image-upload; `stacks/gateway-chat/` (nginx:alpine, ana-docker:8091, homepage tile). (4c40b9f, 740bcae)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **pi on nh3-dev wired to `gen`** — `~/.pi/agent/models.json` + a `gen` launcher mirroring `glm`. Local box config, not a repo artifact. (untracked by operator choice)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **`gen` model → Qwopus3.5-122B-A10B (vision-intact NVFP4), full 256K @ fp8.** ana-ml2 GPU 0, served-name `qwen3.5-122-a10b`. STABLE vLLM + fp8 KV reaches 262144 — nightly+turboquant-4bit-KV was UNNECESSARY. 92.7 tok/s warm. (20e796c, 5b06514)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **TP=2 across the two ana-ml2 Blackwells REJECTED** — `PIX` (PCIe, NO NVLink) → all-reduce-bound → SLOWER. One-model-per-card is optimal; PP=2 only if a >96 GB model is ever wanted. (untracked by operator choice)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **qwen-image-bench (T2I judge, NVFP4) replaced qwen3.6-35b-a3b on GPU 1**, aliased `image-judge`. (bfae924, 5dfce04)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-18]` **heretic abliterated Mistral Small 4 NVFP4 built + LIVE as `mistral-small-4`** (in-house NVFP4, drop-in stack, A/B'd vs official, "heretic stays"). DISPLACED from GPU 0 on 2026-06-19 for the gen/Qwopus swap. (dd3a5c9, f566f61, `tools/mistral-small4-nvfp4/`)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-18]` **irv-ml1 VRAM consolidation + comfy-dev torch-pin** — ComfyUI pinned to the A6000 (48 GB), audio zoo → 3090; `DISABLE_UPGRADES=true` @ torch 2.12.1. (a8550ad)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-18]` **ComfyUI acceleration set (9 nodes) installed for comfy-dev** on irv-ml1.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-17]` **ana-ml2 `vm.overcommit_memory=1` made durable** (sysctl drop-in) — overcommit=0 + zero swap caps CommitLimit; large model-file mmap ENOMEM'd despite ~393 GB free. (fc88eff)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-17]` **GLM 5.2 wired into litellm** (`glm-5.2` + `-reasoning`, z.ai passthrough, `extra_body.thinking.type` toggle). (fe77a35)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-17]` **nh3-extdev stood up as a sudo-LESS infra-ops manager box** (successor to nh3-ansible); hosts `/opt/externs` pi-on-GLM-5.2 client workspaces. (a841eab)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-20]` **Worldtree capability-gateway aliases stood up; NO `wt-` prefix (operator) — reuse the existing capability convention.** gen/gen-reasoning/classifier/summarizer/chat-judge/reranker/scalar-judge (passthrough)/gen-frontier/-reasoning. The **swap-transparency rule**: embedding stays model-named (`qwen3-embedding`) because vectors aren't swap-transparent; everything else can swap behind a stable alias. (d372103, 7819f96)
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **Removed the litellm `*` wildcard → llama-swap** (operator; llama-swap decommissioned, :9292 confirmed down). Unknown/typo'd models now 404 cleanly instead of a misleading "Connection error" — closes the long-standing footgun. (d0eb09c)
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **Langfuse retired** (operator) — redundant with LiteLLM native spend_logs + a 6-container crash-prone surface (the 94 GB clickhouse spew). Callbacks dropped, stack torn down. (f8eda1c)
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **OmniVoice streaming `/tts` — vendor chatterbox-fast's scheduler, not its whole package** (operator-approved vendor-copy). Diffusion TTFA floor finding → tuned 16-step streaming + aggressive packing; chatterbox-fast remains the lowest-latency front-end. (288d085, cd92b85)
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **Daily image+buildx prune cron on ana-docker** (the missing automation behind the recurring image bloat) — `lkraven` crontab, prune unused images >48 h + per-builder buildx cache. The real systemic fix (daemon log cap) is deferred to clean context.
|
||
_Archived 2026-06-25._
|
||
|
||
|
||
- `[2026-06-25]` **althing re-architected to the lean multi-machine bus; nh3-extdev stood up as a
|
||
MODEL B mesh peer** (dedicated `althing-svc` + group-shared `/srv/althing`). `reference_nh3_extdev_althing_mesh`.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-23]` **zellij native web client piloted on nh3-dev** (`zellij-web.service` :8443)
|
||
alongside ttyd. `reference_zellij_web_seat`.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-22]` **Worldtree persona-render config arc (#314/#322/#317) pre-synced + deployed green
|
||
on demo+personal** — #317 a boot-blocking config REMOVAL. `reference_corviduo_dev_emergency_ops`.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **R22 (brokkr/dwarves) stood down to gateway-only; full-access R22 key minted;
|
||
Phase B CANCELLED** (Worldtree model-agnostic → no deploy path). Key at `/home/lkraven/.r22-gateway-key`
|
||
(persistent mode-600, carries paid GLM, don't delete). MUT = free `qwen3.5-122-a10b` (`gen`).
|
||
Operator steer: R22 research is gated on a pragmatic/deployable outcome, not advancing-the-art.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **claude-bot issue-scope token minted for worldtree-dev self-serve** (id 16,
|
||
`write:repository`+`write:issue`); old token revoked. Advances the credential-migration directive.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **rest-server-ana recovered + backup prevention shipped + worldtree-dev admin keys
|
||
provisioned** (demo d113207c / personal f4f75adb). Cred rotation (5 rest-server pw) BELAYED.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **claude-bot → ADMIN on vh/Worldtree** (operator-authorized) — self-serves WT
|
||
deploys/tokens henceforth.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.**
|
||
(auto-memory `project_migrate_infra_access_to_claude_credentials`)
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel provisioned + fix
|
||
verified** (15×→1.01× re-embed). `reference_wt_gateway_scoped_log_view`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-01]` **qwopus native MTP speculative-decode tested on `gen` → NOT kept** (+12% single-stream,
|
||
−15–20% aggregate at concurrency, silently drops min_p/logit_bias). Banked for T1. `reference_gen_qwopus_122b`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-01]` **Deckard trial → reverted to qwopus (`gen`)** (won writing "in every way" but ~36 vs
|
||
~90 tok/s; spec-decode rescue ruled out). git `b63c48b`→`681eb70`. Deckard kept staged as T1's
|
||
writing benchmark.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-02]` **mtf-dev granite harness-spike ran GREEN — MECHANICAL only, efficacy DEFERRED to
|
||
the T1 run.** Trainer TRL SFT→DPO→eval seam proven end-to-end on a synthetic fixture (not the E-RP
|
||
corpus); operator DECIDED no intermediate real-efficacy granite spike (uninterpretable proxy —
|
||
arch gap + abliteration axis). `reference_gen_qwopus_122b`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-03]` **ratatoskr-dev DEMO Heimdall key provisioned (R30 φ0).** Minted a tier-user key on
|
||
the demo via `POST /admin/keys` (bootstrap admin key), mirroring their personal base consumer (no
|
||
character-binding); base-agent affect reads work ungated. `reference_worldtree_demo_key_mint`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-04]` **character-rp: gateway-forwarding proven + role pre-staged + #344 shipped.**
|
||
Empirically confirmed per-request `extra_body` (top_k/repetition_penalty) forwards through the
|
||
`gen-reasoning` LiteLLM alias to vLLM + standard params override the alias defaults — no gateway
|
||
cap needed (I over-built a dedicated alias, operator corrected, reverted with zero fleet impact).
|
||
Pre-staged the `character-rp` role into demo+personal bind-mount `model_roles.yaml` (byte-verified
|
||
on b18; caught the cached-registry ordering). worldtree-dev shipped **#344 (v1.0.0b19)** for the
|
||
durable-agent override-drop. spend_logs spot-check queued (ratatoskr's timestamp ping).
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-04]` **althing v2 herald+receiver formalized as systemd on nh3-dev.** `althing-herald.service`
|
||
(Restart=always, **Environment=PATH incl ~/.cargo/bin** — the pane-dispatch fix) + `althing-receiver.service`
|
||
(v2 → pillar-3 `/owner/*` live); stale forseti unit removed; both on v2.0.0b2, canonicalized by
|
||
forseti. `reference_nh3_dev_althing_herald`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-04]` **LitBench-RM respun (irv-ml1 A6000, comfyui displaced)** for T1's reward ensemble;
|
||
operator sole comfyui consumer, holding image-gen until LitBench done. `reference_litbench_rm_irv_ml1`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-05]` **T1 training venue: CLOUD recommended; operator chose smoke-first on ana-ml2.**
|
||
On-prem ruled out (ana-ml2 full — both 96G GPUs ~93G used): keep-serving = NVMe offload ~6-8 DAYS;
|
||
full ana-ml2 shutdown = CPU offload ~1 DAY but a whole-fleet outage. Cloud Vast.ai 8×80GB (no
|
||
offload → ~3-6h, ~$60-500, zero fleet impact) is the clean alt (mtf-dev + infra-ops both rec;
|
||
Vast for its no-content-AUP marketplace + likely-existing VastBlue account). Operator's next step
|
||
= the ana-ml2 CPU-offload SMOKE (~60 steps) to get real samples/sec before the full-outage-vs-cloud
|
||
call. HF base verified public (zero-auth pull). Runbook + gotchas in `reference_t1_cloud_train_plan`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-05]` **glm-5.2 canonical limits recorded** (probed live vs z.ai): **1,048,576 (1M) input
|
||
context / 131,072 (128K) max output**; NO gateway-side cap (pure passthrough → z.ai's limits are
|
||
effective). Written to the config comment (commit `624a07e`) + `reference_litellm_gateway`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-06]` **infra-ops NOPASSWD sudo confirmed on irv-ml1 + ana-ml2** (the arbo memory's "NOT
|
||
infra-ops on irv-ml1" was stale) — enabled the worker deploy + every ana-ml2 hot-swap smoke.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-07]` **SGLang zc hot-swap: 2-line patch RAN → PARTIAL. Cleared startup + self-attn E2E, GDN
|
||
in_proj hit 2 MORE name-routing gaps → MERGE recommended.** Root cause was a NAME-routing gap (not the
|
||
feared shape-assert): `lora/utils.py:235` hardcoded `params_mapping` lacked `in_proj_qkv/in_proj_z →
|
||
in_proj_qkvz`. Patched those 2 lines (overlay bind-mount) → `init_lora_shapes` passed, server up, self-attn
|
||
loaded end-to-end. But the SAME missing GDN mapping bit 2 MORE resolvers: adapter-load skipped 192 in_proj
|
||
weights, and inference `utils.py:300 get_target_module_name` (substring match) RAISED → 500 + scheduler
|
||
died. So GDN LoRA name-routing is incomplete in ≥3 SGLang paths (fixed 1). NOT shape-intractable (self-attn
|
||
proves the separate-key→fused-buffer machinery works), but NOT "one targeted fix" either → per operator's
|
||
LAST-shot rule → MERGE. mtf-dev looped + closed (thread 01KWZDZFYN…); B/A shard-math never reached.
|
||
Supersedes the "SGLang wants a THIRD format" framing — no new adapter format needed.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-07]` **T1 (retargeted qwopus-122B→AEON-27B) SFT COMPLETE on ana-ml2** — all-TRL torch-GDN,
|
||
batch1/accum8/GC-off, 1 epoch ~9h @ 0.25 s/s, loss 2.445, adapter banked. Supersedes the 2026-07-05
|
||
cloud-vs-smoke venue decision: AEON-27B's 52G bf16 fits one Blackwell → ran on-prem in ~9h (no cloud,
|
||
no full-fleet outage). GC is NOT the throughput lever (batch1-no-GC == batch4+GC at 0.25; dvalin's 0.565
|
||
was warmup-noise from a 30-sample smoke); ~0.25 s/s is the torch-GDN floor (only FLA beats it, sm_120-
|
||
blocked #913). `reference_t1_cloud_train_plan`.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-07]` **T1 hot-swap on vLLM NVFP4 is DEAD — vLLM 0.24.0 qwen3_5 LoRA-application silent no-op**
|
||
(#47639 open / fix PR #47640 open / regression from #37912). Proven quant-agnostic (NVFP4+FP8) AND
|
||
adapter-format-agnostic (separate `zc` loads-but-inert; fused rejected); #47640 guard-patch overlay didn't
|
||
fix. Root-caused by a 3-peer dwarf panel (eitri/regin/dvalin); their vLLM issue refs verified real via
|
||
GitHub API. Engine mirror-image: vLLM wants SEPARATE keys, SGLang wants FUSED.
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-07]` **HARD REQUIREMENT restated: hot-swappable LoRA or ABANDON.** Merge is NOT a solution for
|
||
this requirement (static; reserved for a later RP-quant track). SGLang is the last hot-swap shot; if it
|
||
can't apply, park the adapter + file upstream repros. `reference_lora_training_worker` (adjacent).
|
||
_Archived 2026-07-13._
|
||
|
||
- `[2026-07-07]` **comfy-dev in-arbo LoRA-training worker DONE** — built/deployed/validated Phases 1+2
|
||
(first Sindra train succeeded, publish-step + `train_id`), Phase 2.5 GC contract (shape A) confirmed +
|
||
impl deferred. `services/lora-training-worker/`, commits 888ba6a→b617a8b. `reference_lora_training_worker`.
|
||
_Archived 2026-07-13._
|
||
|
||
## Tried and abandoned (archived)
|
||
|
||
- `[2026-04-30]` task-board workflow with
|
||
`container: image: debian:bookworm-slim` — fails:
|
||
`actions/checkout@v4` needs `node` at runtime, slim image lacks it.
|
||
Switched to `node:20-bookworm-slim` (has node + apt) or runner-label
|
||
default. (Pattern revisited 2026-05-17 for skaldsong-dev: container
|
||
override needs `nodejs` apt-installed unless it IS the default.)
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-04-30]` Dropping the `container:` directive before runner
|
||
re-registration with docker-schema labels — runner silently falls
|
||
back to **host mode** (jobs run inside the alpine `act_runner`
|
||
container itself, no apt). The `:host` suffix in startup logs
|
||
(`labels updated to: [pfi-fleet:host ana-docker:host]`) is the
|
||
giveaway. Fix: register with `pfi-fleet:docker://<image>` schema
|
||
labels.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-04-30]` Updating runner labels by editing `.env` and bouncing
|
||
— doesn't take. The `.runner` registration cache pins labels at
|
||
first registration; env-var updates are read each start but the
|
||
stored token + UUID are tied to the original label set on the gitea
|
||
side. Fix: stop runner, delete `.runner`, generate new admin
|
||
registration token, redeploy.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-04-30]` `git reset --hard origin/<sha>` in
|
||
`deploy-task-board.yaml` (and the in-repo nevermore playbook before
|
||
fix) — invalid syntax: `origin/` prefix only works for branch refs.
|
||
SHAs need `git reset --hard <sha>` directly. Resolved with
|
||
`git rev-parse --verify --quiet "origin/{{ ref }}^{commit}"` first,
|
||
then bare `"{{ ref }}^{commit}"` fallback.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-04-30]` Assuming `DEPLOY_SSH_KEY` was at user scope after
|
||
task-board wiring — it was actually only repo-scope on
|
||
`vh/task-board`. vor's first CI run failed with empty SSH key
|
||
(`printf '%s\n' "" > ~/.ssh/id_ed25519`). Fix: copy secret to user
|
||
scope at `gitea.phasefinal.com/user/settings/actions/secrets`.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-04-30]` `grep -vE "^(#|$)"` to inspect `.env` for sanity —
|
||
leaked the full `MINIFLUX_PASSWORD` line into the transcript. Then a
|
||
follow-up redaction attempt with `sed -E "s/=(.{4}).*$/=\1<redacted>/"`
|
||
still leaked the first 4 chars. Lesson: when probing secret-bearing
|
||
files, use field-by-field SELECTIVE inspection
|
||
(`grep -E "^(KEY1|KEY2)="`) rather than negative filters; for any
|
||
password line, `grep -c` (existence) or `test -n "$(...)"`
|
||
(non-empty), never `cat` or value-printing.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-08]` Filtering Traefik's UTC access log by Gitea-local-PDT
|
||
timestamp substrings (`grep "2026/05/08 15:1[2-7]"`) returned zero
|
||
matches and led to a wrong "no /v2/ traffic in 12 days" conclusion.
|
||
**Gitea logs in PDT, Traefik logs in UTC** — same host, different
|
||
timezones. Always normalize timezones (UTC) when correlating logs
|
||
across services on the same box. Cost: ~30 min in the wrong
|
||
direction.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-08]` Bumping Gitea `PER_WRITE_TIMEOUT` /
|
||
`PER_WRITE_PER_KB_TIMEOUT` to address `unexpected EOF` on
|
||
`/v2/.../blobs/uploads/` PATCH — wrong direction. Both govern
|
||
**response writes**, not request body reads. `unexpected EOF` from
|
||
Go's HTTP server means the client closed mid-body-upload; not a
|
||
knob Gitea exposes server-side.
|
||
_Archived 2026-05-27._
|
||
|
||
- `[2026-05-12]` Defaulting asset-engine to Traefik-routed
|
||
(`asset-engine.phasefinal.com` with `anaprod` cert resolver) on
|
||
first scaffold — user pulled it back to LAN-direct. Internal
|
||
tools default LAN-direct; only add Traefik when an external/TLS
|
||
surface is actually needed.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-12]` Routing althing thread replies through `galdrabok`
|
||
when the actual dev handle is `forseti` — bus rejected
|
||
`to=forseti` initially because thread participants list was
|
||
`[galdrabok, infra]`. Solved by starting a new thread with
|
||
`forseti` as the direct recipient. Lesson: when the bus
|
||
auto-resolves a sender handle that doesn't match the actual dev
|
||
role, start a fresh thread rather than fighting the participant list.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-13]` Initial Voxtral default voice `alloy` (OpenAI-compat
|
||
naming) — vLLM-Omni serving Voxtral does NOT translate aliases.
|
||
Native presets are `<register>_<gender>` shape (`neutral_female`,
|
||
`casual_male`, etc.). Always live-probe `/v1/audio/voices` for the
|
||
exact wrapper-deployed preset names before setting a catalog
|
||
default. Same caveat for Qwen3-TTS (wrapper exposes 15 voices: 9
|
||
Qwen presets + 6 OpenAI aliases) and Kyutai-TTS (NillPointer
|
||
wrapper has NO voice-listing endpoint at all; voices are
|
||
filesystem paths under the `kyutai/tts-voices` HF repo).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` `--task classify` for Skywork in vLLM 0.19.1 —
|
||
flag was deprecated. Use `--runner pooling`; the model's
|
||
`architectures: [LlamaForSequenceClassification]` in config.json
|
||
drives the classification head. Surfaced as `vllm: error:
|
||
unrecognized arguments: --task classify` in container logs.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` Trusting that `.env` edit alone propagates a new env
|
||
var into a worldtree container — `compose.yaml`'s `&worldtree-env`
|
||
anchor must explicitly declare the passthrough or the value silently
|
||
doesn't land. Same footgun bit `Z_AI_API_KEY` (2026-05-12) AND
|
||
`BIFROST_CLIENT_ALLOWED_HOSTS` (2026-05-17). Cost ~10 min of "why is
|
||
env empty?" diagnosis each time. Worldtree-side fix in
|
||
`vh/worldtree@08f02b2`.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-17]` `--force-recreate --pull never` from the docker:cli
|
||
sandbox without explicit `-e WORLDTREE_IMAGE=<sha>` re-pins the
|
||
container to `:latest`, even when a newer SHA-tagged image is on
|
||
disk. Symptom: container "recreated" but actually reverted to a
|
||
stale image. Pass `-e WORLDTREE_IMAGE=...:<sha>` to the docker run
|
||
invocation. Worldtree-dev's `8ef3801` health-gated `:latest`
|
||
advance is the long-term fix.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-18]` Volva env.sh.template `$HOME` in commented examples
|
||
— systemd's `EnvironmentFile=` parser doesn't expand `$HOME`;
|
||
uncommenting lands the literal `$HOME/...` string. Volva-dev's
|
||
`f4dda73` swapped to `/home/<svc-user>/...` placeholders.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-18]` Initial Volva systemd unit's `ProtectHome=read-only`
|
||
without `ReadWritePaths=` — althing-cli's SQLite (`~/.althing/
|
||
althing.db`) and codex's session state (`~/.codex/`) both need to
|
||
write. Container started but every poll failed with "db path not
|
||
writable". Surgical fix: `ReadWritePaths=/home/lkraven/.althing
|
||
/home/lkraven/.codex` (preserves the hardening intent, only carves
|
||
out the specific dirs).
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-18]` Trusting that env.sh's `export VOLVA_ROOT="$(cd
|
||
"$(dirname "${BASH_SOURCE[0]}")" && pwd)"` template line works
|
||
under systemd — `EnvironmentFile=` parser aborts on the first
|
||
unparseable line (command substitution), and `VOLVA_ALTHING_CLI`
|
||
declared below silently never lands. Symptom: `Environment=`
|
||
property empty, daemon error "althing-cli not found at
|
||
'althing-cli'". Fix: replace command-substitution with literal
|
||
path. Volva-dev's `d436c3c` dropped VOLVA_ROOT entirely upstream.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Naive `docker rmi worldtree:<old-sha> --force`
|
||
for CD SHA cleanup — would untag pinned/personal worldtree images
|
||
since all three deployments share corviduo-dev. Use
|
||
`docker image prune -a --filter "until=Xh"` instead — respects
|
||
in-use semantic (Docker won't remove an image referenced by any
|
||
container on the host), so pinned/personal protected automatically.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Skaldsong CD first attempt: `docker pull` step
|
||
failed with 401 unauthorized. ana-docker had no `docker login`
|
||
for `gitea.phasefinal.com`. My playbook prereq note ("docker login
|
||
has been done at least once") was an unverified assumption. One-time
|
||
manual login persists in `~/.docker/config.json`; architectural fix
|
||
(workflow-side `ssh ana-docker 'docker login ...'` step using
|
||
`REGISTRY_USER`/`REGISTRY_TOKEN` secrets) flagged as v2.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` `SKALDSONG_HOST_CORS_ORIGINS=http://10.250.50.70:8300`
|
||
as a bare URL — pydantic-settings parses complex env vars via
|
||
`json.loads()`; first-boot crashloop with `SettingsError: error
|
||
parsing value for field "cors_origins"`. Must be JSON array literal:
|
||
`SKALDSONG_HOST_CORS_ORIGINS=["http://..."]`.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` `SKALDSONG_HOST_STATIC_ASSETS_PATH=/app/web/dist`
|
||
in compose — mismatched Dockerfile reality. The Dockerfile COPYs
|
||
SvelteKit build output flat into `/app/spa` (not `/app/spa/dist`).
|
||
Lifted the path from skaldsong-dev's CD-ask message ("`/app/web/dist`")
|
||
rather than verifying against the actual Dockerfile they shipped.
|
||
Lesson: when encoding container-internal paths in compose, verify
|
||
against the Dockerfile, not the design-doc.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-19]` Playbook verify step `docker ps | grep healthy`
|
||
racing the container's `start_period` (30s in compose's
|
||
healthcheck). Verify ran 0.09s after `compose up -d --force-recreate`
|
||
— well before docker's healthcheck could flip the status from
|
||
`(health: starting)` to `(healthy)`. False-negative; container was
|
||
operationally up (the earlier `/health` poll verify already
|
||
confirmed). Fix: grep `^Up` not `healthy`. /health-200 IS the
|
||
liveness check; docker's `(healthy)` is just a delayed echo.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-20]` `SKALDSONG_DB_PATH` + `SKALDSONG_RUNS_DIR` in
|
||
compose env block — names skaldsong's app doesn't read. App reads
|
||
`SKALDSONG_HOST_SQLITE_PATH` + `SKALDSONG_HOST_RUNS_ROOT` (per
|
||
Dockerfile ENV defaults). Wrong names = silently no-op; app fell
|
||
back to Dockerfile defaults pointing at `/app/data/...` which the
|
||
compose's bind mount did NOT cover (target was `/app/state/...`).
|
||
Result: every `--force-recreate` wiped the SQLite DB. Caught by
|
||
skaldsong-dev (althing thread `01KS4DPF6SXTBP4Q360JZVWPNT`). Fix in
|
||
`52e98fa`. Lesson: verify env var NAMES against the Dockerfile/app,
|
||
not against design-doc shorthand.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-25]` First selene-block patch put the block under
|
||
`saga_allowed_models:` instead of top-level `models:` — used
|
||
`text.replace("models:\n", ...)` which substring-matched the
|
||
`saga_allowed_models:\n` line first. Caused YAML parse error. Fix:
|
||
anchored regex `re.compile(r"^models:\n", re.MULTILINE)`. Pattern:
|
||
substring replace on YAML top-level keys WILL match suffix-containing
|
||
keys.
|
||
_Archived 2026-05-31._
|
||
|
||
- `[2026-05-27]` `docker compose up -d` inside the `docker:cli` sandbox:
|
||
`${VAR:-./config}` defaults resolve `./config` to the sandbox CWD, but the
|
||
Docker daemon interprets the path against the HOST filesystem → auto-creates
|
||
an empty dir → entrypoint reseeded image-baked defaults (lost host-side
|
||
providers.yaml patches). Fix: pass `-e WORLDTREE_CONFIG_DIR=/abs/path`.
|
||
Folded into the docker-as-root convention note.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` `:latest`-pinned compose + private gitea registry + sandboxed
|
||
pull = recreate on ancient cached `:latest` (deploy pulls by SHA so the tag
|
||
never advances; sandbox can't pull). Fix: retag SHA→`:latest` on host, then
|
||
`--pull never`. Better: pin SHA in `.env`, advance in CI.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` Container recreate during in-flight skaldsong gen kills the
|
||
runner. With deploys every ~10min and stories >5min, structural not
|
||
incidental. Roadmap (skaldsong-dev): pre-shutdown signal handler, per-scene
|
||
resume-from-checkpoint, /api/admin/quiesce. None shipped.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` `--reasoning-format gemma` on artemis-31b-v1i — unsupported in
|
||
the deployed llama.cpp (accepts none|deepseek|deepseek-legacy). `deepseek`
|
||
populates the `reasoning_content` SSE delta Worldtree GemmaProvider checks.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-27]` `head -c N` piped after a streaming curl SIGPIPEs the curl,
|
||
killing the request early. Use file-write + separate read.
|
||
_Archived 2026-06-01._
|
||
|
||
- `[2026-05-31]` Building the dia2-capable image surfaced THREE upstream
|
||
packaging quirks: (1) `pip install -e nari-labs/dia2` fails — no PEP 660
|
||
`build_editable` hook; (2) plain `pip install` builds an empty
|
||
`UNKNOWN-0.0.0` wheel (base setuptools 59.6 < dia2's required ≥70); (3)
|
||
`--no-deps` leaves `transformers`/`sphn`/`whisper-timestamped` missing. Fix
|
||
(`local/dia:v2`): copy the pure-python `dia2/` package into site-packages +
|
||
install ONLY those 3 deps; base torch/numpy already satisfy Dia2.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` Dia2 predefined voices (43, baked at `/app/voices`) are NOT
|
||
reachable from the `/tts` clone path — it resolves `clone_reference_filename`
|
||
against the reference_audio dir ONLY. The OpenAI `/v1/audio/speech` `voice`
|
||
param auto-resolves them (separate code path), which masked the gap. Fix:
|
||
stage `/app/voices/*` into `/worktank/dia/reference_audio`. Lesson: verify on
|
||
the endpoint the catalog ACTUALLY targets.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` `voice_mode=clone` with an empty `clone_reference_filename`
|
||
→ asset-engine serializes it as the literal string `"undefined"` → `/tts`
|
||
404. First observed on dia2; worked around in the catalog (default the field
|
||
to a real voice). **[2026-06-01] root cause found** — the Kokoro voice-blend
|
||
widget reading Shoelace `.value` before hydration (see Current state); the
|
||
real fix is asset-engine-side and is escalated.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` asset-engine `services.schema.json` is DERIVED (regen from the
|
||
Pydantic model via `dump_schema.py`) and had DRIFTED — rejected the
|
||
`lifecycle` field 12/14 services use. RESOLVED: regenerated with
|
||
`CatalogLifecycle` + `reproducibility.seed_field` (`b7b2130`). Lesson:
|
||
hand-editing `services.yaml` shape without regenerating re-introduces drift.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-31]` ttyd-over-TLS forces HTTP/2 (kills ttyd's terminal WebSocket
|
||
→ blank screen); Safari/WebKit never sends HTTP basic-auth on WS upgrades.
|
||
Both solved for the fleet seat (Caddy forces HTTP/1.1; auth → network-gating)
|
||
— detail in auto-memory `reference_ttyd_fleet_seat`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-05-30]` esh-docker-vm: `hard` NFS mounts from 10.0.50.50 froze a
|
||
container worker in UNKILLABLE D-state when the NAS stalled — only a host
|
||
reboot clears it. Separately, `fstab defaults` (no `_netdev`) made NFS-bind
|
||
containers `Exited(255)` on reboot. → auto-memory
|
||
`incident_esh_docker_nfs_boot_race`.
|
||
_Archived 2026-06-03._
|
||
|
||
- `[2026-06-02]` **Naive cumulative re-decode for streaming** — decode the
|
||
growing token prefix each chunk + emit the delta, assuming the causal flow
|
||
gives a stable prefix. It does NOT (full-context attention, see above);
|
||
maxdiff 0.30. Also the high-level `s3gen.inference` re-applies `trim_fade` at
|
||
the start every call. Don't go this way.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **CSM bring-up** — upstream `phildougherty/sesame_csm_openai`
|
||
Dockerfile pins NO `huggingface_hub` version, so it now resolves to 1.17.0
|
||
where `huggingface-cli` is fully removed (replaced by `hf`) → the build dies
|
||
at `huggingface-cli login`/`download`. Passing the HF token as a build-arg
|
||
got the token in (the model downloads at BUILD time, not runtime — our
|
||
compose only wired it as runtime env, the original failure) but then hit the
|
||
rotted CLI. Fixing would need vendoring + patching the Dockerfile; not worth
|
||
it for a weak model → CSM removed.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Fish path-form references** (`{audio:"/app/references/x.wav"}`)
|
||
→ HTTP 500 on this build, despite the stack README showing that shape. The
|
||
working clone paths are `reference_id` (by name → resolves the DIRECTORY
|
||
`references/<name>/` containing `<name>.wav` + `<name>.lab`; the flat
|
||
`references/<name>.wav` is ignored — see the two-root-causes entry above) or
|
||
inline base64 `references`. The catalog uses `reference_id`.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Reference transcript barely affects Fish clone IDENTITY**
|
||
(disproving my mid-session theory). I'd blamed a bogus provenance-note `.txt`
|
||
for poor cloning, but the ECAPA re-test showed correct (0.787) / bogus (0.778)
|
||
/ empty (0.738) transcripts all clone Imogen about equally — the transcript
|
||
affects PRONUNCIATION (phoneme disambiguation per the docs), not who it sounds
|
||
like. The real culprit for "not British" was the `"undefined"` select bug, not
|
||
the transcript. (A correct transcript still marginally helps pronunciation —
|
||
cheap to stage, not load-bearing.)
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **Pitch-shift register control** (rubberband, to deepen Imogen
|
||
to contralto/mezzo) — Fish ignores small reference shifts and overshoots
|
||
large ones; `-3st` WITHOUT formant preservation → "baritone through a tuba"
|
||
(formants collapsed); `-2/-3st` WITH `--formant` → too subtle / still bad.
|
||
Abandoned at every depth; all variants deleted. Finer independent
|
||
pitch/formant control needs praat (not installed). Future clones = operator's.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-01]` **resemblyzer is too weak for cloning-fidelity scoring** — its
|
||
dated 2019 LSTM encoder + a synthetic-vs-natural domain gap scored the Imogen
|
||
clone CLOSER to the default than to real-Imogen, which led me to a WRONG "Fish
|
||
is a weak cloner" call. ECAPA-TDNN (speechbrain `spkrec-ecapa-voxceleb`) on the
|
||
same clips gave the correct answer (clone 0.79 to real Imogen, 0.10 to
|
||
default). Use ECAPA, not resemblyzer, for speaker-verification.
|
||
_Archived 2026-06-04._
|
||
|
||
- `[2026-06-02]` **Fish (fish-s2 / OpenAudio S1-mini) progressive streaming — SHELVED
|
||
(sub-realtime).** Benched RTF on A6000: 0.72x/0.82x/0.86x, mean 0.80x = sub-realtime,
|
||
so client-side chunking starves. Root cause: Fish only chunks on `<|speaker:X|>` tags;
|
||
plain text → batches=[whole text] → all semantic tokens generate before any audio
|
||
(chunk_length inert). Plus kui/ASGI StreamResponse doesn't flush (fix = anti-buffering
|
||
headers, kept on file). STANDING REVISIT TRIGGER: when an RTX Blackwell Pro lands → bench
|
||
fp4-quantized Fish; if RTF > ~1.5x give it the chatterbox-fast treatment. For now Fish is
|
||
a buffered catalog entry (great for SAVED gens, not live-audition).
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-02]` **Context-priming at chunk joins (chatterbox-fast §1.6) — ABANDONED
|
||
(discard-cut leaks the prefix).** Prepend prior sentence, generate `prefix+content`,
|
||
discard prefix audio. Built + shipped (`d707439`), live-A/B'd, reverted (`090e70a`).
|
||
Kill: `generate()` returns one waveform with NO marker where the prefix ends, and the
|
||
model renders the same prefix with different timing solo vs followed-by-content — so
|
||
locating the cut is a guess that left a whole clause in the output (operator caught
|
||
"...without a trace of sarcasm" spoken twice). Reliable cut needs token-level boundaries
|
||
or per-chunk forced-alignment (heavy). Coherence loss at joins = ACCEPTED limitation.
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-02]` **Native frame-level streaming on Chatterbox-TURBO — ABANDONED (turbo
|
||
isn't built for streaming).** (1) `S3GenStreamer` is docstring-only (unimplemented).
|
||
(2) `flow_inference(finalize=False)` is buggy (lookahead trim removes frames from `h`
|
||
not `h_masks`/conds → shape mismatch); a 1-line patch makes it RUN. (3) BUT the flow
|
||
encoder uses full-context attention (`static_chunk_size=0`) → incremental decode is
|
||
prefix-unstable (adding tokens re-attends, shifts earlier mel ~0.30-0.39 vs one-shot).
|
||
(4) Forcing `static_chunk_size>0` didn't stabilize it. True sub-second frame-streaming
|
||
needs deep attention surgery with quality risk — not worth it. → adaptive-chunking instead.
|
||
_Archived 2026-06-11._
|
||
|
||
- `[2026-06-03]` **`uv pip install .` fails on SmoothKen/knn-svc** (and similar script-repos)
|
||
— it's analysis scripts + a poetry pyproject, no buildable package (setuptools
|
||
package-discovery error). Install the pyproject deps directly, don't build the "package".
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **Backgrounding `althing-cli monitor` with an inline shell `&` (instead of
|
||
the Bash-tool `run_in_background`) orphans it** — it survives the shell exit, holds the
|
||
per-handle flock UNTRACKED (won't notify the session), and `stop-monitor` doesn't detect it.
|
||
Fix: find + kill the orphan PID (verify cwd=this repo / handle first — nh3-dev is shared, other
|
||
agents' monitors run there too), then re-arm via run_in_background. Always re-arm tracked.
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **torch-2.12 venvs need `uv pip install torchcodec`** — torchaudio 2.12
|
||
defaults to the TorchCodec backend for `.load`; without it, real audio I/O throws "TorchCodec
|
||
is required" — and it ONLY surfaces at actual conversion, NOT at import/model-load. Lesson:
|
||
validate real I/O, not just import, when provisioning ML engine envs. (seed-vc on torch 2.4
|
||
uses the legacy backend, exempt.)
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-03]` **gitea webhook to a private IP is denied by `webhook.ALLOWED_HOST_LIST`**
|
||
(anti-SSRF; default `external` blocks private/loopback). Symptom: delivery shows
|
||
`dial tcp ...: webhook can only call allowed HTTP servers`. Fix = APPEND the target net to
|
||
ALLOWED_HOST_LIST in gitea's app.ini (keep `external`; scope tight, never `*`/`private`) +
|
||
restart gitea (act_runner job containers survive a restart). gitea runs as a container on
|
||
ana-docker (`gitea_gitea_data` volume, `/data/gitea/conf/app.ini`).
|
||
_Archived 2026-06-13._
|
||
|
||
- `[2026-06-04]` **homepage "crash" ≠ always NFS** — a wedged container in unkillable D-state can come from dead `siteMonitor` widget targets (retired ESH firewall IPs) hanging the node event loop into `exit_mmap`, needing a host reboot. Check homepage's siteMonitors against retired hosts. (`incident_esh_docker_nfs_boot_race`)
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **pyworld needs `setuptools<81`** (imports the removed `pkg_resources`); and **R/soundgen `-lgfortran` fails** on irv-ml1 (default gcc is gcc-11 but only gfortran-12 present) → install `libgfortran-11-dev`.
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **Persistent services on irv-ml1 need SYSTEM systemd units** — the box reaps user-session processes on ssh disconnect, and `--user` systemd isn't reachable over non-login ssh, so nohup/setsid/`screen -dmS`/`systemd-run --user` all die (even with enable-linger). Use `/etc/systemd/system/`.
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **GPU pin by INDEX is ambiguous on irv-ml1** — native CUDA orders fastest-first (A6000=0) but nvidia-smi/docker use PCI order (A6000=1). Pin by **UUID** (`CUDA_VISIBLE_DEVICES=GPU-…`); verify via nvidia-smi compute-apps.
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-04]` **Ollama/llama.cpp-BUNDLED chat templates silently diverge from canonical HF — the "headgun" lesson.** Serve each model's canonical `tokenizer.apply_chat_template`, not the bundled template (bundled ones corrupt baselines). Verify via vLLM `/tokenize`→`/detokenize`. (`90e08f0`/`27eb537`)
|
||
_Archived 2026-06-14._
|
||
|
||
- `[2026-06-08]` **Demucs `uv pip install demucs` pulls torch 2.12/torchaudio 2.11 → `ta.save()` requires torchcodec → dies AFTER separating** (0 stems, rc=1). Fix = pin `torch==torchaudio==2.4.1` + `UV_LINK_MODE=copy`. Validate the SAVE path, not just import + GPU inference, on a bleeding-edge torch.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **vLLM 0.19 CUDA-graph-capture OOMs on a SHARED GPU** — fills KV to the `--gpu-memory-utilization` budget WITHOUT reserving graph-capture memory, so `capture_model` OOMs AFTER weights+KV load (crash-loops). Fix: free co-tenant room OR `--enforce-eager`. FP8 single-stream is batch-1 GEMV (memory-bound) → Q4 wins single-stream by physics; FP8 wins under concurrency. (`reference_ana_ml2_vllm_granite`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Langfuse has NO public dashboard-creation API** — dashboards/widgets are postgres rows; clone a default + swap the measure. tok/s is the `outputTokensPerSecond` MEASURE (metrics-API/dashboard query time), not a per-generation field. langfuse-web needs `HOSTNAME=0.0.0.0`. Host 3000 is gitea's → langfuse on 3001.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **`sudo` over non-interactive ssh FAILS SILENTLY where the user lacks NOPASSWD** (esh + corviduo are OUTSIDE the infra-ops identity) → empty output misread as "empty file." Read world-readable files WITHOUT sudo. corviduo ssh = `vh@10.250.50.152`; bind-mount configs are vh-owned, the `.env` is deploy-owned 600.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-05]` **Worldtree summarizer-model is NOT an env var** — no `WORLDTREE_SUMMARIZER_MODEL`; defaults to claude-haiku in code, opt-in via config not `.env`. Inspect the live container env + vh-owned config files first.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Loading an old LoRA adapter after a vocab bump fails on embedding size.** The harness config + `tokenizer.json` are now at `new_vocab_size=50279` (v4 multi-tag); the v1/v2/v3 adapters are 50277. To load v1 (the accepted artifact), set `cfg.new_vocab_size=50277` before `load_finetuned_engine_lora` (else PeftModel state_dict size mismatch). `tokenizer.json.v3bak` is the 50277 tokenizer for a clean restore.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Heavy from-source compile (`MAX_JOBS=128`) on the shared PROD GPU box PINS it** — load hit 187, prod vLLM restarted, killed an in-flight quant. ana-ml2 hosts live inference; never run a big build there at full parallelism. Cap `MAX_JOBS≤32`, build off-box, or cgroup-constrain.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **`--quantization fp8` on a VL model can quantize the VISION TOWER → garbage vision** (Qwen3.5-VL on stable vLLM: gray-grid output; LM answers text fine, so it "looks" healthy). The nightly excludes the vision tower. Validate the VISION path on a quantized VLM, not just text — and pin the engine digest with the exclusion.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **vLLM's `--gpu-memory-utilization` is checked against FREE VRAM at startup, not total** — on a shared card, growing one service before trimming a co-tenant OOMs. Trim the shrinking service FIRST, then grow. Size to the FREE budget.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **The `vllm/vllm-openai` entrypoint is already `["vllm","serve"]`** — compose `command:` supplies the model as the first POSITIONAL arg + flags; a second `serve`/`--model X` → "unrecognized arguments". Same-class: `tee` masks the real exit code (use `>`); HF `datasets` rejects bare `wikitext` (needs `Salesforce/wikitext`).
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-13]` **Chatterbox-Turbo LoRA finetune: the repo's `setup.py` loads the WRONG tokenizer** — pulls gpt2-medium + a grapheme merge file instead of the chatterbox-turbo GPT2 tokenizer (vocab.json+merges.txt, len 50276). Fix = override + delete the grapheme `tokenizer.json`; `[vmoan]` → new_vocab_size 50277 (1-row resize), lora_r 64 / alpha 128, modules_to_save=[text_emb,text_head]. Unique-stem corpus collision needs `{index}_{stem}` IDs. (`irv-ml1:~/r16-vmoan-harness`)
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-11]` **A completion-poll `while pgrep -f <scriptname>` SELF-MATCHES its own remote shell argv** — its own `pgrep -f` always finds itself → the loop never exits. Use a match pattern ABSENT from the poll command (the python stage, or a sentinel file), not the driver's own name.
|
||
_Archived 2026-06-16._
|
||
|
||
- `[2026-06-14]` **vLLM ModelOpt-NVFP4-MoE loader is broken for current multimodal MoEs.** `nvidia/Qwen3.6-35B-A3B-NVFP4` fails weight-load: `KeyError: layers.0.mlp.experts.w2_input_scale` on 0.19.1, `lm_head.input_scale not registered` on 0.22.0 (vllm #44081); same class hits Gemma-4 MoE / Qwen3-30B-A3B NVFP4. The arch + quant ARE recognized (gets past arch resolution + vision-processor load) — it's the per-expert/lm_head scale-key mapping. Don't chase nightlies; use official FP8 until fixed.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **vLLM sampler-warmup OOMs on a shared GPU even when weights fit** — it warms the sampler with `max_num_seqs` (default **1024**) dummy requests, and a big vocab (Qwen3.6 = 248K) makes that a huge transient logits tensor. A vision endpoint doesn't need 1024-way concurrency: set `--max-num-seqs 32`. Separately, post-load `ValueError: No available memory for the cache blocks` means util is too thin (weights+activation+graph ate it) — for 34 GB FP8 weights, util ≥ ~0.45 to leave KV room.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **Recreating multiple vLLM services concurrently races the memory-profiling assertion** — `AssertionError: Error in memory profiling. Initial free memory X / current Y … other processes … release GPU memory while vLLM is profiling`. Recreate co-tenant vLLM services ONE AT A TIME (force-recreate one, wait healthy, next).
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **embed/rerank (0.6B) at util 0.05 reserve ~5.5 GB each — mostly util-reservation WASTE, not need.** A 0.6B model needs ~1.2 GB weights + ~2.5 GB CUDA/torch context; util 0.03 (~3.6 GB) fits with room, reclaiming ~4 GB (vLLM reserves the util fraction regardless of actual KV; embedding models barely use KV). Real-need floor ~3 GB — don't go to 0.02.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-14]` **Chatterbox-Turbo decode-knob foot-guns** (R16 v1-polish + emergent probes): the turbo length cap is `max_gen_len` (default 1000) on `t3.inference_turbo`, NOT `max_new_tokens` — and `tts_turbo.generate` does NOT forward it (wrap inference_turbo to cap). `rep_pen 2.0 / temp 0.5` BACKFIRES (degenerate 24 s run-on). Soft decode tamers cut the NVV ITSELF, not just the run-on tail (operator: "p1 trims the moaning too") — gen-time polish can't beat v1's defaults. Inline base-NVV tokens DEGENERATE (moan-cascade + gibberish), they don't modulate the surrounding words.
|
||
_Archived 2026-06-18._
|
||
|
||
- `[2026-06-15]` **A comfyui restart-loop with `exit 0` + `OOMKilled=false` is NOT OOM — suspect a native SEGFAULT.** `docker inspect` shows only the LAST exit code; a GPU OOM is a catchable `RuntimeError: CUDA out of memory`, NEVER `Fatal Python error: Segmentation fault`. Diagnostic: `RestartCount>0` = same container restarted in-place → the docker json-log ACCUMULATES across all cycles, so grep the full `docker logs` for the crash signature + the lines right before each restart banner. comfyui's was at VAE decode (sampling completes, then segfault loading `AutoencodingEngine`) — gen-triggered but native, not VRAM. Suspects logged: torch/cu130 mismatch, an injected `aimdo` cuda-funchooks lib, the VAE.
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **NEVER blanket-kill `althing-light-monitor` by PID** — they're per-handle and MANY agents run their own on the shared nh3-dev box (heid / ratatoskr-dev / worldtree-dev / bifrost-dev each have one). A `kill $(pgrep -f light-monitor)` downs four peers' monitors (near-miss caught by checking `/proc/<pid>/environ` `ALTHING_HANDLE` + cwd first). Recover your OWN via `althing-cli stop-monitor` (v0.14.1+); a raw `&`-backgrounded monitor is untracked (no harness fire-notification) — always arm via `run_in_background`. (extends the pkill-self-match lesson below)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **Manual `docker compose up` on corviduo Worldtree DOWNGRADES to a stale local `:latest` → crash-loop (config/image mismatch).** Compose `image:` is `…/worldtree:latest`, but local `:latest` lagged ~2wk (`21c90d5eeb2d`); CI/CD deploys NEWER SHA images, and a real pull 401s (daemon has no gitea-registry creds — that's CI/CD's). `--pull never` then used the stale latest → personal crash-looped on `ConfigurationError: echo.default_model missing` → **~2-3 min personal outage**. RECOVER: retag `:latest` to the image a HEALTHY SIBLING runs (`docker inspect worldtree-worldtree-api-1 --format '{{.Image}}'` → `docker tag …`), then `up -d --pull never`; verify `Application startup complete`. Don't manually recreate Worldtree on corviduo without pinning the correct image — image+config are CI/CD-coupled. (auto-memory `reference_corviduo_dev_emergency_ops`)
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **`.claude.json` (Claude Code config) truncates on ENOSPC → invalid JSON → CC resets to defaults on next start.** nh3-dev hit 100% disk (a peer's R19 dataset pull staged ~147G to local /home instead of the 7TB NAS); the in-flight `.claude.json` write was cut mid-key. Repair = strip the dangling token + close the open braces, validate `json.load`. CC keeps its OWN rolling backups at `~/.claude/backups/` + writes a `.claude.json.corrupted.*` snapshot when it self-detects. After ANY ENOSPC on a box running CC/agents: integrity-check `.claude.json` (json.load) AND the althing bus DB (`PRAGMA integrity_check`).
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-15]` **`pkill -f althing-light-monitor` SELF-MATCHES the killing shell** (the pattern is in the command's own argv) → kills itself mid-run (exit 144/truncated output). Stop the light-monitor via `althing-cli stop-monitor` or a captured PID — never `pkill -f <string-present-in-this-command>`. The singleton lock can also RACE to 2 live monitors during re-arm churn; keep exactly one tracked (run_in_background) monitor, and a raw `&` monitor is untracked (no harness fire-notification — don't use it).
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-14]` **Fleet/colo hosts must reach gitea over the INTERNAL route, NOT the public IP.** `gitea.phasefinal.com` = public `38.120.12.44` (ana-srv1); gitea is a container on ana-docker, git-SSH `10.250.50.70:222` + HTTP `:3000`. A fleet host egressing to public `:22` gets fail2ban-banned after any retrying git loop → silently wedges webhook auto-deploys (`git fetch` times out under `set -euo pipefail`, aborts before reset). Bit irv-ml1's arbo deploy. `:22` on `10.250.50.70` is ana-docker's HOST sshd (deploy key → Permission denied), NOT gitea. Documented `docs/orientation.md` (`6e58e57`).
|
||
_Archived 2026-06-19._
|
||
|
||
- `[2026-06-16]` **litellm 500 `Router.acompletion()/aembedding() missing 'messages'/'input'` = a request missing `Content-Type: application/json`, NOT a gateway outage.** curl `-d` defaults to form-encoding → litellm can't parse the JSON body → `data` reaches the router without `messages`/`input` → 500 (should be a 400; litellm #16993). My own diagnostic calls dropped the header → I misread it as a gateway outage and needlessly bounced the gateway ~4× chasing a phantom (image/version/config were fine throughout; a malformed UI-added "Mistral Story Eval" model in the DB was a red herring I deleted). ALWAYS send `-H "Content-Type: application/json"` testing litellm; reproduce with a header'd call before declaring a litellm incident.
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **LiteLLM admin-UI playground can't test vLLM-backed models** — it auto-sends empty `tools:[]`, vLLM 400s (litellm #6228); the gateway `strip_empty_tools` hook is a PROXY hook and structurally can't reach the UI's in-process `litellm.completion()` call. Off-ramp = `tools/gateway-chat.html`. (Langfuse playground also out: its SSRF guard blocks internal-IP LLM connections, wontfix Langfuse #13097.) (auto-memory `reference_litellm_ui_playground_vllm_deadend`)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **A gitea Actions run that "looks like it never fired" (no image, untouched build dir) usually DID fire then skipped or failed-fast.** Bit arbo TWICE: run #3 looked-never-fired but SKIPPED (change-detection `engine=0`); run #4 looked-never-fired but FAILED in ~7s (early-step). Check the run LIST + status via claude-bot (`GET /repos/{o}/{r}/actions/tasks`) BEFORE a runner-connection/dispatch theory. CORRECTION (2026-06-16, proven): action logs ARE readable via API at the per-JOB endpoint — `GET /api/v1/repos/{o}/{r}/actions/runs/{run}/jobs` → take a job `.id` → `GET /api/v1/repos/{o}/{r}/actions/jobs/{job_id}/logs` returns the full plain-text log (claude-bot basic-auth, internal route `10.250.50.70:3000`). Only the RUN-level `/runs/{id}/logs` 404s; the per-job one works, so no UI needed. Gotcha when reading: gitea marks **every** step `conclusion:failure` once any step fails — the per-step conclusions mislead; trust the log body + step *timestamps* (the step with real duration is the culprit) not the conclusions. (Used 2026-06-16 to diagnose arbo run #5/task 1175 = test-gate `uv run pytest` → `Failed to spawn: pytest` = pytest absent from the fresh-checkout venv sync, NOT the SSH-clone comfy-dev suspected.)
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **vLLM 0.23.0 breaks Mistral MULTIMODAL at startup** (`MistralCommonImageProcessor
|
||
has no attribute fetch_images`, #44911 regression ~0.22.1+; same class hit Mistral-3.1/
|
||
Devstral). The dummy-image profiling crashes. Workarounds: `--limit-mm-per-prompt
|
||
'{"image":0,"video":0}'` loads text-only; for VISION, pin **v0.22.0** (last pre-regression).
|
||
Tradeoff: v0.22.0 *buffers* reasoning (doesn't stream reasoning_content cleanly); a newer
|
||
build streams it but breaks vision — vision won (Mistral is the abliteration-baseline).
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **Mistral `-reasoning` TTFT variance = Triton kernel JIT-compile-during-
|
||
inference, NOT reasoning time.** `jit_monitor` logs the spike (`_compute_slot_mapping_kernel`,
|
||
`apply_token_bitmask_inplace_kernel` ← the tool-call/structured-output path). vLLM's startup
|
||
warmup covers only SOME shapes; a first-hit uncovered shape JITs a kernel (seconds), then
|
||
CACHES it — warm until container recreate (the Triton cache is IN-container, no volume, so
|
||
recreate = cold again). Secondary: GPU idle-down to P8/180 MHz between requests. Also:
|
||
Mistral `reasoning_effort` is BINARY (none|high only — medium/low → HTTP 400).
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **Worldtree old-baseline instances need the FULL config set synced, not the
|
||
changeset delta.** Personal (a PRE-mistral baseline) crash-looped (41 restarts) on
|
||
`DanglingBindingError: role 'echo' binds missing catalog entry 'qwen3.6-35-a3b-heretic'` —
|
||
its `model_roles.yaml` (2026-06-10, never migrated) still bound the retired heretic that the
|
||
synced v0.35.10 providers removed. The "model_roles unchanged — skip" was correct only for
|
||
demo (already at the v0.35.9 baseline). Old-baseline → sync providers + model_roles + defaults.
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **vh is a Gitea USER, not an org → no scoped package-write for a service
|
||
account.** Gitea user-namespace packages are owner/admin-only, so claude-bot can't push to
|
||
`vh/arbo` without being admin. The fix is converting vh → a `corviduo` ORG (tabled). Also:
|
||
vh's `tea` token lacks `write:admin` — create gitea users via the server-side `gitea admin
|
||
user create` CLI (`docker exec -u git gitea …`), not the admin API.
|
||
_Archived 2026-06-20._
|
||
|
||
- `[2026-06-16]` **`docker compose up -d` on a multi-service vLLM stack recreates ALL its
|
||
services** (shared `.env` → all configs re-hash) → concurrent memory-profiling race
|
||
("No available memory for the cache blocks"). Target one: `docker compose up -d <service>`.
|
||
(auto-memory `feedback_compose_up_recreates_whole_stack`)
|
||
|
||
_76 older entries archived to archival-memory.md._
|
||
_Archived 2026-06-20._
|
||
|
||
|
||
- `[2026-06-19]` **litellm CONFIG-loaded models can't be hot-removed — `/model/delete` is DB-only.** Only a process restart drops a config model. The INVERSE works: `/model/new` live-ADDS (no bounce) but dups with config on restart. Clean no-bounce-add: live-add via API, delete the DB copy before the eventual bounce.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **the litellm `*` wildcard routed to DECOMMISSIONED llama-swap → a stale/typo'd model name returned a misleading `Connection error`, not `model not found`.** Bit a brokkr call to the renamed `qwen-image-judge`. **→ RESOLVED 2026-06-20: wildcard removed (d0eb09c) — unknown models now 404 cleanly.**
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **A MoE card can't be filled to 0 bytes free — the FusedMoE transient workspace (~3.09 GB) is the floor**, allocated OUTSIDE `gpu-memory-utilization`. FIX: `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` + leave ~3.2 GB free (util <= ~0.95 on a tight MoE).
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **vLLM decode tok/s: ALWAYS discard the first generation (cold warmup).** Cold ~ 24.8 tok/s vs warm 92.7 — a cold read undersells decode ~3-4x. Measure run 2+ over a >=256-token output.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-19]` **For full native 256K on one 96 GB card, nightly+turboquant-4bit-KV was unnecessary for the Qwopus MoE.** Stable fp8 KV already fits 262144 (hybrid SSM+attn caches KV only on attn layers → small pool). Reach for fp8 first.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-18]` **mmartial `comfyui-nvidia-docker`: root pip installs CRASH-LOOP the container.** `docker exec -u 0 pip install` leaves root-owned files in the uid-1000 venv → the boot script (uid 1000) fails its torch step → crash loop. FIX: `chown -R 1000:1000 /comfy/mnt/venv`; pin with `DISABLE_UPGRADES=true`. (auto-memory `reference_irv_ml1_comfyui_mmartial`)
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-18]` **ComfyUI custom nodes break on version-assumption drift** — box runs 0.24.1; TeaCache imports `precompute_freqs_cis` at MODULE level (guard the LTX-only import). `pip install tensorrt` defaults to cu13 → use `tensorrt-cu12`.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-17]` **Mistral HF→NVFP4 quant: the placement knob is the whole game.** `device_map="cpu"` (CPU-resident, sequential onload) works; `"auto"` OOMs, `max_memory` offloads to meta. Read shards with plain `read()` + `safetensors.torch.load(bytes)`, NOT `safe_open` (mmaps the 50 GB shard → ENOMEM on `/tank` ZFS).
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-17]` **HF-format Mistral Small 4 is UNSERVEABLE on vLLM** — no HF `Mistral4` backbone; serves ONLY via the native loader. A HF-format quant MUST be converted to native first.
|
||
_Archived 2026-06-21._
|
||
|
||
- `[2026-06-20]` **A full disk wedges gitea's Actions notifier SILENTLY — and it STAYS wedged after disk recovers; only a `docker restart gitea` clears it.** After the 94 GB clickhouse disk-full event, gitea stopped CREATING workflow runs on push (push succeeds, ref lands, but NO run spawns — newest run stale, queue empty), while the act_runner stayed registered. NOT run-pickup, NOT disk-now (52% free), NO gitea errors/panics — a silent notifier stall. Hit BOTH vh/arbo (comfy-dev) and vh/Worldtree (worldtree-dev's deploys) → it's global, not repo-specific. Fix: `docker restart gitea` on ana-docker (it had been up 2 weeks, never bounced). A restart does NOT retroactively fire pushes made during the outage — re-push/re-trigger after. Reinforces the deferred docker-daemon-default log-cap fix (the disk-full that started it can't be allowed to recur).
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **Mistral-Medium-3.5-128B NVFP4 (RecViking) on ana-ml2 GPU 0 — ABANDONED, too slow (operator confirmed).** Stood up as a temporary speed-check displacing mistral-small-4; measured base **~19.2 tok/s** decode on one Blackwell (vs `gen`/Qwopus3.5-122B ~92.7 tok/s warm) → confirmed the operator's "too slow" suspicion. KEEP these serving findings: (1) vLLM **nightly** loads the HF-format NVFP4 weights DIRECTLY — no Mistral native-convert (unlike Small 4, same `mistral3` arch) — via the FlashInfer Cutlass NVFP4 kernel + `turboquant_4bit_nc` KV; TP=1 fits the ~70 GB NVFP4 on one 96 GB Blackwell. (2) **EAGLE was a dead end**: the official native FP8 head (`mistralai/Mistral-Medium-3.5-128B-EAGLE`) crashes in CUDA-graph replay with graphs on, and nets NO speedup with `--enforce-eager` (~18.5 vs ~19.2 base) — likely low draft acceptance against the NVFP4 (non-official-FP8) base. Untracked workspace files (`stacks/mistral-medium-3.5/` + `playbooks/serve-mistral-medium-3.5.yaml`) removed. ⚠️ Weights still staged on ana-ml2 `/tank/aimodels/mistral-medium-3.5-{nvfp4,eagle}` (~70 GB+) — pending operator OK to free.
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **A 94 GB unrotated `langfuse-clickhouse` json.log was the REAL disk-filler (not image bloat) — there was NO log rotation anywhere.** No per-stack `logging:` limit and no docker-daemon default → ClickHouse's verbose stdout grew unbounded to ~94 GB and filled ana-docker's `/` (healthchecks then failed fleet-wide on `runc /tmp: no space left`). Lesson: on a full disk, the symptom (all-unhealthy) is the FORK-can't-write healthcheck, not the containers; find the runaway via per-container-dir `du`, not `docker system df`. The systemic fix is a daemon-default `log-opts` (`max-size`/`max-file`) — NOT just per-service patching.
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **OmniVoice's upstream "RTF 0.025 / 40×" does NOT hold — diffusion has a ~fixed per-call overhead → TTFA floor + multi-chunk streaming is a net loss.** Measured ~2.8× (32-step) / 5.6× (16-step) on the 3090; every chunk costs ~the same regardless of length, so naive per-sentence streaming triples total gen time + starves. Use few/aggressive chunks; chatterbox-fast (autoregressive) stays the low-TTFA front-end.
|
||
_Archived 2026-06-25._
|
||
|
||
- `[2026-06-20]` **Named buildx builders accumulate cache the standard prunes MISS.** `docker image prune -a` + `docker builder prune` don't touch `buildx_buildkit_<name>-builder0_state` volumes (per-builder buildkit state). Prune them with `docker buildx prune --builder <name> -af`, or `docker volume rm` orphaned ones (builder already removed).
|
||
_Archived 2026-06-25._
|
||
|
||
|
||
- `[2026-06-25]` **althing "unreachable: <machine>" can MASK an app-level 500.** Raw network was
|
||
clean; root cause = receiver DB agents-table not synced with the config roster → delivery 500'd
|
||
"unknown to: <handle>", MAPPED to "unreachable". Diagnose: raw curl to :8087 + connect-probe ⇒ NOT
|
||
network. Fixed in althing v0.17.1. `reference_nh3_extdev_althing_mesh`.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **rest-server `.htpasswd: permission denied` = the ana-nas NFS mount FAILED (ghost
|
||
file on the local mount point), NOT a decommission.** `mnt-backup.mount` stuck `failed` (fstab bare
|
||
`defaults`) → rest-server serves an empty local dir. Recovery in disaster-recovery.md.
|
||
_Archived 2026-07-08._
|
||
|
||
- `[2026-06-20]` **The DEFAULT `ssh ana-docker` is `lkraven` (no NOPASSWD) — but `ssh
|
||
infra-ops@ana-docker` HAS NOPASSWD root.** A `sudo cp` as lkraven silently failed → nearly punted
|
||
the rest-server recovery. Reach for `infra-ops@ana-docker` for sudo ops.
|
||
_Archived 2026-07-08._
|
||
## Archived 2026-08-02 — Recent decisions (archived)
|
||
|
||
### 2026-07-08-worldtree-mimir-deploy-blocker-resolved-mid-session
|
||
|
||
- `[2026-07-08]` **worldtree Mimir deploy-blocker resolved (mid-session):** synced `persona.envelopes.assistant` +
|
||
`envelope_grants:[]` verbatim from the baked canonical into BOTH corviduo-dev instances (demo+personal),
|
||
YAML-validated via each container's own parser; worldtree-dev cleared to push the Mimir-bound image. (Was my
|
||
parked R32 1C envelope-mirror come due — see [[reference_corviduo_dev_emergency_ops]] config-sync recipe.)
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-two-parked-items-closed-phantom-qwen3-6-35b
|
||
|
||
- `[2026-07-09]` **Two parked items closed: phantom `qwen3.6-35b-a3b` alias VERIFIED already-gone; ana-docker
|
||
docker log-cap SOLVED no-bounce.** (1) **Phantom**: absent from `/v1/models` + `/model/info` (config+DB
|
||
registry), zero litellm log refs — the parked "400s in /v1/models" note was STALE (already cleaned in the
|
||
2026-07-08 gen repoint to `-heretic`); bare token survives only in 2 config COMMENTS (lines 76/80). Nothing to
|
||
remove. (2) **Log-cap**: running containers were UNCAPPED (182M json-logs, top offender 59M) because
|
||
daemon.json's `max-size 10m/max-file 3` only applies to containers CREATED AFTER a daemon restart — it never
|
||
reaches already-running ones. No-bounce fix = `/etc/logrotate.d/docker-containers` (**copytruncate** — dockerd
|
||
opens json-logs `O_APPEND` so truncate-in-place resets cleanly, no sparse-file corruption; `size 10M`,
|
||
`rotate 3`, `compress`, `su root root`), auto-picked-up by the daily `logrotate.timer`. Force-ran + gzipped the
|
||
frozen `.1` archives → **182M → ~55M** (44M active + 11M gz), every container kept its multi-week uptime
|
||
(zero bounce, verified). **LATENT FOOTGUN FLAGGED (not yet fixed, operator's call): daemon.json declares
|
||
`live-restore:true` but the RUNNING daemon has it FALSE** (daemon.json was edited after the last daemon start,
|
||
never reloaded) → the NEXT `systemctl restart docker` / crash / pkg-upgrade **bounces ALL ana-docker containers
|
||
once**. Fix WITHOUT a bounce = `systemctl reload docker` (SIGHUP loads live-restore into the running daemon;
|
||
log-opts are NOT SIGHUP-reloadable, which is why logrotate — not the daemon cap — is the enforcer for running
|
||
containers).
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-morpheus-tts-off-the-shelf-voice-pipeline-shipped
|
||
|
||
- `[2026-07-09]` **mOrpheus TTS off-the-shelf voice pipeline SHIPPED end-to-end (irv-ml1) + wired into
|
||
gateway-chat.** Full arc (commits): gen served-name honesty rename aeon→qwen3.6-35b-a3b-heretic (99a4a17,
|
||
vLLM served-name + litellm refs, so /v1/models + spend-logs name the real model); permanent 2-container stack
|
||
(01eedd8); gateway-chat auto-voice quoted dialogue (c948013); streaming decode TTFA 4.5s→0.8s (da76829);
|
||
max_tokens 1200→2400→3500 with a context-clamp (f363fe6, 0655a37 — long lines were clipping at 14.6s, and
|
||
`repetition_penalty` 1.1 is LOAD-BEARING: at 1.0 the model never stops); AudioContext resume-on-gesture
|
||
no-sound fix (033f368); pre-chunk by QUOTED SECTION not sentence for prosody (a1f3023→f295cc1); staged clone
|
||
voices baddy/beatrice/whisper (0655a37 + runtime .wav/.txt in the voices dir); agent voicing prompt (a573514).
|
||
**Load-bearing config, all encoded in stacks/mOrpheus/: bf16 not FP8, image v0.23.0 not latest, GPU=3090 not
|
||
A6000, rep_penalty 1.1.** Serving-viability confirmed: vLLM concurrency near-linear to 8× (707 tok/s).
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-granite-gen-memory-extractor-bind-host-synced-on
|
||
|
||
- `[2026-07-09]` **granite→gen `memory_extractor` bind host-synced on demo+personal Worldtree (Vuong-directed,
|
||
#335 Slice-4).** Changed `model_roles.yaml` memory_extractor `binds.catalog_id` `summarizer`→`gen` (overrides
|
||
intact: thinking:false/temp0/8192) on BOTH `/opt/worldtree{,-personal}/config`; `memory_distiller` left on
|
||
`summarizer` (range-scoped sed `/memory_extractor:/,/memory_distiller:/` — the naive global replace would've hit
|
||
both); backups `*.bak-preqgen-20260709`; validated via each container's OWN yaml parser. **DEMO LIVE on gen**
|
||
(b43 `d501e516732d` auto-deployed mid-edit + its restart RACED my edit by ~2min → I restarted
|
||
`worldtree-worldtree-api-1` to activate; healthy, live process resolves memory_extractor=gen). **PERSONAL
|
||
NOW LIVE on gen too** — Vuong authorized the restart (via wt-dev); restarted `worldtree-personal-worldtree-api-1`
|
||
after a PRE-FLIGHT that ran the app's OWN `load_model_roles()` (`core/llm/roles.py:121`) against the synced config
|
||
INSIDE the running `c9986cd` container: `gen` in catalog, all 9 roles resolve, no `DanglingBindingError` → proven
|
||
safe on the OLDER image BEFORE touching it (model_roles-delta-alone clean; no full-config-set sync needed).
|
||
StartedAt 20:50:55Z, healthy, resolves gen. **BOTH instances live on gen.** **LESSON:
|
||
the bind-mount `/opt/worldtree*/config` SHADOWS the baked `/app/config-defaults/` → the deploy alone never
|
||
updates the ACTIVE config; the host edit is required AND a restart activates it (role registry cached at boot) —
|
||
pre-stage BEFORE the deploy's restart or you race it.** FOLLOW-UPS (non-blocking): (a) `memory.extractor.user_pass`
|
||
parity block → self-serve from the b43 baked `defaults.yaml` (pydantic-default no-op); (b) stale `gen` provider
|
||
description (Mistral-Small-4 → qwen3.6-35b-a3b-heretic) → wt-dev owns the REPO-side `providers.yaml` fix
|
||
(operator's call — NOT purely cosmetic: the gen/dialogue + classifier entries carry Mistral-Small-4 SAMPLING
|
||
defaults that drive mask/lofn/forseti/mimir dialogue, so wt-dev re-validates for qwen), host cosmetic sync pairs
|
||
when it deploys. Gave wt-dev the VERIFIED canonical 4-alias set (backends+samplers read from the LIVE gateway
|
||
config, not the doc); corrected `docs/pfi/model-sampler-defaults.md` seat 4 (had lagged QwQ-RpR-v4 → Deckard-PKD;
|
||
live gateway was always Deckard). Operator SCOPED IN the character-RP re-point (2026-07-09):
|
||
character→char-rp / thoughtful-character→char-rp-reasoning (character-rp per wt-dev's role semantics), moving
|
||
character RP off the GENERAL qwen onto the dedicated Magidonia/Deckard seats. Relayed to wt-dev w/ the mapping
|
||
principle + a SAMPLER-OVERRIDE warning (DROP character-rp's old temp0.75/top_p0.85 overrides — carried onto the
|
||
dedicated seats they'd clobber the canonical RP tuning DOWNWARD) + ratatoskr-reach note (role call is transparent
|
||
but Magidonia/Deckard quality/latency differs from gen). DONE 2026-07-09: wt-dev committed 5d4fa4a (v1.0.0b44,
|
||
UNPUSHED — operator drives push); operator directed host-ahead-of-push, so I sourced BOTH config files directly
|
||
from that unpushed commit (local `~/development/Worldtree` checkout — capital W; `git show 5d4fa4a:config/…`) +
|
||
mirrored VERBATIM to `/opt/worldtree{,-personal}/config` on demo+personal, paired-pre-flighted via the app's
|
||
`load_model_roles()` (no DanglingBinding), restarted both → LIVE: character→char-rp, thoughtful-character +
|
||
character-rp→char-rp-reasoning, memory_extractor→gen preserved. Backups `*.bak-prerp-20260709`. context_window
|
||
VERIFIED (llama.cpp /props + char-rp-gguf `.env`): char-rp **131072**, char-rp-reasoning **262144** (gave wt-dev
|
||
to patch the repo from its interim 32768). **HOST AHEAD of repo-remote until the operator pushes 5d4fa4a** (baked
|
||
config converges with the bind-mount on push+deploy). GOTCHA: demo≠personal — PERSONAL was already partly
|
||
re-pointed (2026-07-06 AEON-era character→char-rp) so its delta was mostly stale-AEON-descriptions→Magidonia/Deckard
|
||
+ character-rp + stripping personal's char-rp `default_params` temp0.7/top_p0.8 that CLOBBERED the gateway RP tuning
|
||
downward; DEMO had no char-rp catalog entries at all (b44 adds them). Diffed each instance vs b44 before applying
|
||
(both deltas = expected changeset only, nothing instance-specific clobbered). wt-dev PATCHED the context_window in **b45/3384a37**
|
||
(char-rp 131072, char-rp-reasoning 262144). BUMPED HOST-AHEAD on both instances (operator-directed 2026-07-09):
|
||
mirrored b45's providers.yaml → `/opt/worldtree{,-personal}/config`, restarted, verified LIVE (char-rp ctx
|
||
131072, char-rp-reasoning 262144, bindings intact); backups `.bak-b44interim-20260709`. HOST now = **b45
|
||
canonical** (providers.yaml) + b44 (model_roles unchanged b44→b45). STILL PENDING: (1) operator's batched push
|
||
of **b44+b45** (`5d4fa4a`+`3384a37`) to converge the repo-remote — host is ahead, no fork; (2) user_pass parity
|
||
block (defaults.yaml — NOT in either, separate). Threads `01KX3SGH`/`01KX48QP` (worldtree-dev),
|
||
monitor armed. See [[reference_corviduo_dev_emergency_ops]].
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-granite-gen-memory-extractor-bind-green-lit-for
|
||
|
||
- `[2026-07-09]` **granite→gen memory_extractor bind GREEN-lit for worldtree-dev (Worldtree #335 Slice 4).**
|
||
Answered their VRAM/concurrency headroom check: gen (qwen 35B-A3B heretic) has ample headroom for ~2 bursty
|
||
idle-triggered extractor calls (fixed 0.37 util; KV pool runs 0-2%; --max-num-seqs 16, near-linear batching).
|
||
Corrected their stale "gen = Mistral Small 4 / 119B-6B" belief (gen IS the qwen 35B-A3B heretic since 2026-07-08).
|
||
**This bind is INDEPENDENT of the full granite RETIRE** (reclaim ~32GB on ana-ml2 GPU1) — that stays the
|
||
operator's call, pending brokkr R33 portfolio + production-concurrency due-diligence. Tracked: althing thread
|
||
01KX3SGH… (worldtree-dev) + brokkr's gen-absorbs-granite consult (thread 01KX2V32…) + [[reference_litellm_gateway]].
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
## Archived 2026-08-02 — Tried and abandoned (archived)
|
||
|
||
### 2026-07-01-a-personal-worldtree-ci-deploy-that-fails-85s
|
||
|
||
- `[2026-07-01]` **A personal-Worldtree CI deploy that fails ~85s in with "not found / unauthorized"
|
||
is usually the pull-only-vs-build RACE, not registry-auth.** `deploy-personal.yml` is PULL-ONLY but
|
||
fires on the `staging/vX` tag simultaneously with `deploy.yml`'s build → pulls before the push
|
||
finishes. FIX: re-run once built, or gate on `workflow_run: completed`.
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-04-on-prem-t1-train-that-keeps-any-ana
|
||
|
||
- `[2026-07-04]` **On-prem T1 train that keeps ANY ana-ml2 serving up = ~6-8 DAYS** (1-GPU + NVMe
|
||
ZeRO-Infinity offload; MoE ~10B-active cuts FLOPs but NOT the 244G base's param I/O). The only fast
|
||
on-prem path is a FULL ana-ml2 shutdown (both GPUs + the ~421G vLLM RAM freed → base fits in the
|
||
566G CPU RAM) → CPU offload → ~1-day full-fleet outage. Cloud (no offload) = hours. `reference_t1_cloud_train_plan`.
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-04-a-systemd-user-daemon-that-shells-out-to
|
||
|
||
- `[2026-07-04]` **A systemd `--user` daemon that shells out to `~/.cargo/bin`/`~/.local/bin` tools
|
||
needs an explicit `Environment=PATH`** — the minimal `--user` default silently drops them. The
|
||
althing herald lost `zellij` → silent `pane-miss` for ALL config-backed TUI/pane agents; CC + FIFO
|
||
routes were unaffected, so it was invisible from a CC session. `reference_nh3_dev_althing_herald`.
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-fp8-breaks-morpheus-audio-token-generation
|
||
|
||
- `[2026-07-09]` **FP8 breaks mOrpheus audio-token generation.** `--quantization fp8` on the 3B → 0 valid SNAC
|
||
frames even at GREEDY (degenerate audio+text mix, no start-of-speech); bf16 is clean (28/28 frames). Quant-breaks-
|
||
TTS, same class as the Angel-NVFP4 lesson below. bf16 is REQUIRED (so the operator's "util 0.1" wish is moot — the
|
||
bf16 weights alone are 6.6GB). NB the raw-token benchmark RTF 0.50 was fp8+graphs = never real.
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-sentence-chunking-tts-loses-prosody
|
||
|
||
- `[2026-07-09]` **Sentence-chunking TTS loses prosody** — generating each sentence cold flattens the intonation that
|
||
spans a line. Chunk by QUOTED SECTION (whole quote = one gen call). Also: `repetition_penalty` >1.1 BREAKS cloning
|
||
(penalizes the ~1100 in-context reference audio tokens; keep ≤1.1 on the clone path).
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
### 2026-07-09-hf-whisper-datasets-aren-t-actually-whispered
|
||
|
||
- `[2026-07-09]` **HF whisper datasets aren't actually whispered.** Claris-Whispered-English measures voiced 0.8
|
||
(not a whisper) + IPA transcripts; `datasets` audio decode needs torchcodec (wants CUDA-13, incompatible w/ the
|
||
cu124 venv). LPC DSP-whisperize went unstable (NaN). **kokoro `af_nicole` IS a genuine whisper** (voiced 0.24) —
|
||
that (operator's pointer) is the working whisper reference source, not TTS-voice screening or dataset-hunting.
|
||
|
||
_Archived 2026-08-02._
|
||
|
||
|
||
## Recent decisions (archived 2026-08-03 batch)
|
||
|
||
### 2026-07-14-nvfp4-mtp-fast-char-rp-reasoning-seat-landed
|
||
|
||
- `[2026-07-14]` **NVFP4+MTP fast char-rp-reasoning seat LANDED + LIVE + gateway-repointed + VRAM-tuned.** Modelopt-format re-quant made MTP work. The load-crash root cause = **vLLM 0.24 does NOT propagate modelopt `exclude_modules` to the spec-decode DRAFT model** → the bf16 mtp head gets quantized → shape crash; NO checkpoint config fixes it (`is_layer_skipped` is exact-membership, and the drafter never sees the mtp excludes) → **workaround = a mounted `sitecustomize` that force-skips `mtp.*` in `is_layer_skipped`** (upstream vLLM bug to file). Productionized as compose stack `heretic2-charrp-reasoning` (:8018, workaround baked in). Gateway `char-rp-reasoning` alias fixed: repointed off the stale GGUF served-name `deckard-pkd-27b`, added `enable_thinking:true`, **dropped `min_p`** (MTP-incompatible), canonical samplers temp1.0/top_p0.95/top_k20. Rebalanced GPU0 (gen 0.37→0.30/16-seq/256K + reasoning 0.39/16-seq/192K+MTP + char-rp 128K, 2.7GB free). All 4 gateway roles verified; vLLM reasoning-parser confirmed **leak-free** (unlike the GGUF budget-forcing). Full record + the 4 quant landmines in `docs/runbooks/heretic2-nvfp4-mtp-seat.md`; committed `982c319`. Open (non-blocking): brokkr P00 (seat is live ahead of it), retire the stopped GGUF reasoning seat, file the vLLM bug.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-14-nvfp4-quant-chase-resolved-gibberish-pivoted-to-modelopt
|
||
|
||
- `[2026-07-14]` **NVFP4 quant chase RESOLVED (gibberish) + PIVOTED to modelopt for MTP.** One ~40-min GPU0 window. Root-caused the `!!!!` to the quant NAMESPACE (text-only `AutoModelForCausalLM`→`model.layers.*` keys; vLLM serves only `Qwen3_5ForConditionalGeneration`, which needs `model.language_model.*`) — found from config diffs + vLLM source with ZERO GPU time; fixed by loading as `AutoModelForImageTextToText`. NVFP4 now serves COHERENT (validated greedy). BUT base NVFP4 ≈53 tok/s ≈ GGUF's 59.5 at batch-1 (no single-stream win) AND MTP = 0% acceptance on compressed-tensors (bf16 mtp head only loads on the modelopt format). Operator chose to **pursue a modelopt-format re-quant** (the only path to the 2-4× MTP goal; AEON-proven on this exact Qwen3.6-27B arch). Scoped + de-risked: AEON `/tank/aimodels/qwen36-27b-aeon-nvfp4` = the modelopt reference (quant_method modelopt, 1967 tensors, 15 bf16 mtp keys identical to graft); nvidia-modelopt 0.45.0 installs + `mtq.quantize`/`NVFP4_DEFAULT_CFG`/`export_hf_checkpoint` API confirmed; pipeline unchanged except swap llm-compressor→modelopt. Seats restored; char-rp-reasoning stays GGUF. Full plan in Current state ★ section.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-14-pursue-the-nvfp4-mtp-fast-char-rp-reasoning
|
||
|
||
- `[2026-07-14]` **Pursue the NVFP4+MTP fast char-rp-reasoning seat to completion** (Vuong-directed via /snapshot: "chase the nvfp4 quant, we know it works, write down the recipe"). Full recipe + diagnostic ladder in Current state / in-flight above. Artifacts on ana-ml2 `/tank/aimodels/heretic2-nvfp4-work/` + scripts committed in eshpfi `services/heretic2-nvfp4-quant/`.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-14-char-rp-reasoning-seat-deckard-pkd-neo-code
|
||
|
||
- `[2026-07-14]` **char-rp-reasoning seat: Deckard-PKD → NEO-CODE = Heretic2-Thinking (Qwen3.6-27B)** — R36 gate PASSED (tools 0.967, #355 runaway ELIMINATED). #355 was MODEL-level (Deckard emitted qwen3_coder XML malformed → mangled args → retry-runaway), NOT the reasoning-budget bug; NEO-CODE emits it clean. Custom llama.cpp KEPT (qwen3_coder parse — stock b8840 predates it — + PR#25544). Committed f960a73; full record auto-memory [[charrp-custom-llamacpp-pr25544]].
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-14-soong-lab-webhook-auto-deploy-real-root-cause
|
||
|
||
- `[2026-07-14]` **soong-lab webhook auto-deploy real root cause = gitea `webhook.ALLOWED_HOST_LIST`** (was `external, 10.100.0.0/16` = NH3-only; blocked corviduo-dev's Anaheim `10.250.x` → gitea refused to deliver, never opened the connection). Fixed to fleet-wide `10.0.0.0/8` (app.ini `[webhook]`) + gitea restart; listener now logs every delivery. The ufw `10/8` open (also this session) was a real-but-secondary gap. Committed 462d528.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-355-residual-root-cause-supersedes-the-litellm-gateway
|
||
|
||
- `[2026-07-13]` **#355-residual ROOT CAUSE (supersedes the "LiteLLM gateway holds while seat idles" entry below — that was DISPROVEN).** char-rp-reasoning enters a non-terminating REASONING loop (tool-call-retry planning) and runs to `max_tokens=32768` (~22 min @ 24.7 tok/s, ~13% of requests); the seat GENERATES all 32768 tokens (not idle), and `--reasoning-budget 400` is NOT enforced. 3-source-confirmed (spend_logs completion_tokens=32768 ×4; seat eval-time log; pcap 100%-`reasoning_content` deltas). Server-side fix wanted (operator: no max_tokens ceiling) → routed to brokkr (accepted, pulled dvalin). Lesson (again): confirm before concluding — the seat-idle claim came from reading only the ≤73s requests + missing the concurrent 32768-token slots. See ACTIVE 1.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-deploy-speed-real-bottleneck-uv-sync-memory-s
|
||
|
||
- `[2026-07-13]` **Deploy-speed real bottleneck ≠ uv sync (memory's assumption was wrong).** Buildx step log: `chown -R /app` = 251s (copy-up of the root-owned venv into a fresh layer), uv sync only 35.6s, registry layer cache already wired. Fix = drop `/app` from the chown (validated safe: zero /app runtime writes on both live instances) + uv cache-mount. Shipped as PR #359 (branch off origin/main@b60), worldtree-dev green-lit. Expected ~5min off (~11→~6min). Runner-side BuildKit cache task (b) was already done → moot.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-wt-355-residual-300s-hang-localized-to-our
|
||
|
||
- `[2026-07-13]` WT #355 residual 300s hang localized to OUR LiteLLM gateway (holds 2 char-rp-reasoning requests ~21 min while the seat idles), NOT the seat — Deckard seat EXONERATED (completes ≤72s; `--reasoning-budget 400` forecloses a mid-thinking hang). Corrects worldtree-dev's "seat wedging" diagnosis. Decisive next = the FIN-check (pcap on corviuo). See in-flight ACTIVE 1. **[SUPERSEDED 2026-07-13 — see the ROOT CAUSE entry above; the gateway-hold/slot-leak theory was disproven, the seat was generating 32768 tokens.]**
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-wt-355-turn-lifecycle-fix-validated-on-worldtree
|
||
|
||
- `[2026-07-13]` WT #355 turn-lifecycle fix VALIDATED on worldtree b60 — wedged turns self-terminate cancelled/stalled at the 300s stall-watchdog (turns 2064/2065 vs pre-b60 2061's 16-min no-terminal). worldtree-dev filed follow-ons #356 (rehydrate Tier-3 ctx on resume — the recreate-durability gap), #357 (reclaim orphaned active-turn locks), #358 (LLM-provider read-timeout audit); surfacing to Vuong to prioritize.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-worldtree-deploy-bottleneck-the-image-build-11-min
|
||
|
||
- `[2026-07-13]` Worldtree deploy bottleneck = the image build (~11 min of a ~12 min deploy), root cause the Dockerfile `uv sync ... --no-cache` + no BuildKit cache-mount (re-downloads all deps cold every build). Fix split: worldtree-dev Dockerfile cache-mount diff + infra-ops runner-side persistent BuildKit cache. Config-only changes skip the build entirely (pinned recreate).
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-13-ledger-tier-3-consumer-ledger-miranda-provisioned-on
|
||
|
||
- `[2026-07-13]` Ledger tier-3 consumer `ledger:miranda` provisioned on personal :8081 (key b38932f5, GPG-delivered+shredded, allowlist 10.100.10.50:8770 live); `assistant`+`thoughtful-assistant` capability roles added (gen/gen-reasoning) on personal+demo, canonical d8bd497. Chosen instance = personal (the tier-3-consumer instance, ratatoskr+soong-lab colocated).
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-10-heimdall-grant-ratatoskr-affect-full-on-personal-worldtree
|
||
|
||
- `[2026-07-10]` **Heimdall grant: ratatoskr `affect.full` on PERSONAL Worldtree (operator-approved, worldtree-dev
|
||
R34-v1 request).** Added allow-rule `ratatoskr-affect-full-allow` to `/opt/worldtree-personal/config/policies.yaml`
|
||
(`principal.user_ids:["ratatoskr"]`, action `affect.full`, resource `*`, effect allow), mirroring the #347
|
||
`session-history-write-ratatoskr` rule exactly + placed right after it. **WHY user_ids-based (not tier):** ratatoskr's
|
||
personal key is the minimal **readonly-admin** observability tier, which is NOT in the tier-based
|
||
`affect-render-baseline-allow` (anonymous/user/free/pro/admin) → needs an explicit user_id grant, same as #347.
|
||
R34-v1 (b46, committed UNPUSHED) gates `affect.emit` `dominant_emotion` egress by exposure ceiling (affect.full|safe
|
||
→ present; neither → null); this grant keeps ratatoskr's view alive across the b46 deploy. Surgical exact-string
|
||
insert (preserves comments), backup `policies.yaml.bak-pre-affectfull-20260710`, validated via the CONTAINER's own
|
||
yaml parser (35 rules, +1, payload confirmed). **NOT restarted — deliberate:** rule is on the bind-mount (shadows
|
||
baked), INERT until b46 gating ships, so the b46 CI/CD deploy restart activates it (no live-session blip now). Demo
|
||
untouched (personal-only per key scope). Replied to wt-dev (thread 01KX6DB3…) offering an immediate restart if they
|
||
want it live for pre-b46 testing. See [[reference_corviduo_dev_emergency_ops]].
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-10-comfyui-v0-27-1-success-on-irv-ml1
|
||
|
||
- `[2026-07-10]` **ComfyUI v0.27.1 SUCCESS on irv-ml1 (operator-confirmed execute-now) — landed on torch 2.12.1,
|
||
SageAttention preserved, crash-loop AVOIDED.** The prior attempt (entry below) crash-looped because a torch
|
||
2.12.1→2.13 bump broke SageAttention's ABI. This time I checked `git diff v0.24.1 v0.27.1 -- requirements.txt`
|
||
FIRST and found **core v0.27.1 leaves `torch` UNPINNED** → the version bump does NOT require torch 2.13 (that came
|
||
only from the mmartial boot-upgrade). So: `git checkout v0.27.1` (clean tree) → `pip install -r requirements.txt`
|
||
as **uid 1000** with a **torch-pin constraint file** (torch/vision/audio pinned to current +cu129) to block any
|
||
transitive bump → torch stayed 2.12.1, SageAttention 2.2.0 untouched. Added decord 0.6.0 (fixed SAM3Segment).
|
||
`docker restart comfyui` → healthy, `/system_stats` comfyui_version=0.27.1, "Using sage attention", HTTP 200, DB
|
||
migrated 0003→0004. Reported the divergence to comfy-dev (thread 01KX6D3C…, reply pending) + asked whether LTXV 2.3
|
||
needs a separate torch-2.13 follow-up (their domain; Krea-2's ≥0.25 need is met by 0.27.1). **LESSON: before a
|
||
mmartial ComfyUI version bump, `git diff <old> <new> -- requirements.txt` — if torch is unpinned, bump the CODE
|
||
without touching torch (constraint-pin it) and compiled exts (SageAttention) survive. `docker exec` lands as uid
|
||
1025(comfytoo), not 1000 — use `-u 1000` + the venv python `/comfy/mnt/venv/bin/python`.** See
|
||
[[reference_irv_ml1_comfyui_mmartial]].
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-10-comfyui-0-25-x-bump-on-irv-ml1
|
||
|
||
- `[2026-07-10]` **ComfyUI 0.25.x bump on irv-ml1 ATTEMPTED → FAILED → ROLLED BACK (snapshot saved it).** comfy-dev
|
||
requested (Vuong-authorized) bumping the irv-ml1 `comfyui` stack (mmartial image, `/opt/docker/compose/comfyui/`,
|
||
0.24.1) to 0.25.x for Krea-2 + LTXV 2.3. **TWO FINDINGS: (1) `DISABLE_UPGRADES=false`/USE_PIPUPGRADE bumps the
|
||
VENV (torch 2.12.1→2.13.0 + deps) but does NOT advance the ComfyUI CODE checkout** (`/comfy/mnt/ComfyUI` =
|
||
`/worktank/comfyui/run/ComfyUI` stayed 0.24.1 — pinned/detached git, comfy-dev's domain). **(2) the torch bump
|
||
broke SageAttention** (2.2.0 `_fused.so` undefined-symbol `c10::impl::cow::materialize_cow_storage` vs torch
|
||
2.13.0) → `--use-sage-attention` (REQUIRED launch flag in COMFY_CMDLINE_EXTRA) crash-looped ComfyUI. Net: broke
|
||
the working state, zero 0.25.x payoff. **ROLLBACK WORKED**: pre-bump 16G venv snapshot
|
||
`/worktank/comfyui/venv-snapshot-comfyui-0.24.1-20260710.tar` restored (torch 2.12.1 + working SageAttention),
|
||
re-pinned DISABLE_UPGRADES=true, recreated → healthy on 0.24.1, serving :8188. Broken venv parked at
|
||
`/worktank/comfyui/run/venv.broken-torch213-20260710`. **CORRECTED PATH (sent comfy-dev, thread 01KX655V):**
|
||
comfy-dev git-advances the ComfyUI checkout to 0.25.x + reqs → I handle the torch bump + SageAttention
|
||
rebuild-against-2.13.0 + re-pin (snapshot stays as the net). **LESSON: mmartial `DISABLE_UPGRADES` gates ONLY
|
||
the venv pip-upgrades, NOT the ComfyUI git checkout; a torch bump breaks compiled exts (SageAttention) →
|
||
rebuild-after is mandatory.** Bump BLOCKED pending comfy-dev's git-advance. Stack: A6000 (NVIDIA_VISIBLE_DEVICES=1),
|
||
lkraven-owned compose+venv (uid 1000, no sudo needed), COMFY_CMDLINE_EXTRA OOM flags preserved.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-10-biweekly-open-weight-releases-scan-cron-set-up
|
||
|
||
- `[2026-07-10]` **Biweekly open-weight-releases scan cron set up for brokkr-smithy (Vuong-authorized).** Durable
|
||
systemd **--user** timer on nh3-dev (`brokkr-landscape-scan.timer`, OnCalendar `*-*-01,15 09:00:00`
|
||
America/Los_Angeles, Persistent=true; linger on) → `.service` → wrapper `~/.local/bin/brokkr-landscape-scan.sh`
|
||
runs headless `claude -p "$(cat ~/.config/brokkr-landscape-scan/prompt.txt)" --dangerously-skip-permissions` in
|
||
`~/development/brokkr-smithy` (ALTHING_HANDLE=brokkr-smithy-dev; **explicit PATH** — the --user minimal-PATH
|
||
footgun; per-run logs `~/.local/state/brokkr-landscape-scan/`). Prompt = brokkr's payload verbatim (LLM/image/TTS
|
||
new-release sweep → ranked synthesis → commit+push+notify). VALIDATED: git-push non-interactive (BatchMode
|
||
ls-remote to gitea, passphraseless key — no agent), headless claude auth (READY smoke). VALIDATED END-TO-END 2026-07-10 (manual run, exit 0):
|
||
web-sweep→synthesis→commit `2ed2f29`→PUSH of scan #2 (open-weight-releases-2026-07-24.md); triaged dwarf input +
|
||
caught baseline errors, quality strong. **HANDLE-COLLISION caught+FIXED** — the headless scan shared handle
|
||
brokkr-smithy-dev with the LIVE session + raced its inbox (eitri's dwarf-reply got stolen by the live monitor);
|
||
registered a dedicated **brokkr-scan-dev** handle (`add-handle`, driver=none) + repointed the wrapper + rewired
|
||
step-5 notify → `althing-cli post --to brokkr-smithy-dev` (NO vuong althing handle exists — confirmed). model=default
|
||
+ `--max-turns 80`. First run under the new handle = 7/15. Off-cycle 07-24 doc is a validation artifact (scheduled
|
||
1st/15th runs date to their own run-date, no collision) — operator naming-convention call pending.
|
||
**NEXT AUTO-RUN 2026-07-15 09:00 PDT.** Manual validation/first run = `systemctl --user start
|
||
brokkr-landscape-scan.service`. Open w/ brokkr (thread 01KX63G6): confirm notify-Vuong handle/mechanism + session
|
||
handle + model/turn-cap. **NEXT brokkr task (operator-sequenced after this): TTS audition env** — Higgs-TTS-3 +
|
||
ZONOS2 + Chatterbox baseline, TTFA/RTF + blind-A/B web-listen (thread 01KX6371; needs GPU-placement + HF-token
|
||
feasibility pass first; brokkr delivers the prompt set after the env's up; protocol doc in brokkr-smithy repo).
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-08-rp-seat-campaign-closed-char-rp-magidonia-24b
|
||
|
||
- `[2026-07-08]` **RP-SEAT CAMPAIGN CLOSED — char-rp = Magidonia-24B-v4.3 (128K), char-rp-reasoning = Deckard-PKD
|
||
Qwen3.5-27B (256K); both GGUF/llama.cpp on ana-ml2 GPU0 alongside gen (35B-A3B, util 0.37), ~4G GPU0 margin.**
|
||
Arc: (1) replaced broken Angel NVFP4 with Magidonia prose + QwQ-RpR-v4 reasoning (b268f93); (2) max-context via q8_0
|
||
KV (f570604); (3) canonical samplers for all 4 gateway seats, dvalin-derived + char-rp A/B-tuned (aac4bcf);
|
||
(4) rebalanced gen 0.40→0.37 to fund char-rp 128K (f49c4e4); (5) RE-A/B'd the reasoning seat (operator wanted a
|
||
DRY-tolerant model): **Deckard WON** on brokkr's frozen scorer (composite 2.176, 0/30 loops, 0/30 refusals) over
|
||
RpR-v4 (3.716, 1/30 loop), Pantheon-Reasoning (1.383 but 7/30 refusals), Snowdrop+Gembrain (llama.cpp
|
||
template-incompat) — deployed (5f79b40); (6) Deckard→256K (41305bf); (7) dvalin CONFIRMED Deckard samplers = the
|
||
live A/B set is canonical (4954ca0). **GATE LESSON: a llama.cpp reasoning seat needs a STOCK template that natively
|
||
opens `<think>`/`enable_thinking` (Qwen3.x/QwQ pass; ChatML + Gemma-4 fail) — no monkeypatching. INFRA: llama-swap
|
||
b8840 can't load Qwen3.6/Gemma-4 archs → `ghcr.io/ggml-org/llama.cpp:server-cuda` (pulled on ana-ml2).**
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-08-off-the-shelf-inference-pivot-executed-serve-curated
|
||
|
||
- `[2026-07-08]` **OFF-THE-SHELF INFERENCE PIVOT executed — serve curated abliterated models, stop home-training.**
|
||
Final topology: **gen = `llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-NVFP4-Experts-Only`** (LIVE, modelopt, vision,
|
||
util 0.40), **char-rp = an RP unicorn to be found on fresh context** (see Current state). Intermediate steps
|
||
ABANDONED: Pantheon-Reasoning-27B (served briefly as gen — refuses dark fiction via DeepSeek-distilled
|
||
refusal-reasoning, see Tried); Pantheon-27B-with-MTP for RP (bf16 MTP won't load on the compressed-tensors path);
|
||
Angel MS3.2-24B (my NVFP4 quant = garbage). Prefer EXISTING community NVFP4/GGUF quants over self-quanting
|
||
("don't quant unless you have to" — operator). GGUF serving is now on the table for RP (NEVER Ollama). Gateway
|
||
sampling-defaults wiring still PENDING.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-08-dpo-was-silently-running-3-epochs-harness-gap
|
||
|
||
- `[2026-07-08]` **DPO was silently running 3 epochs (harness gap) → KILLED at epoch 1.2, retargeted to 0.3
|
||
epochs (operator call).** Root cause: `DpoConfig` had NO `epochs` field + `_dpo_config_kwargs` didn't pass
|
||
`num_train_epochs` → DPO fell through to trl DPOConfig's default 3.0 (SFT correctly pins 1 via SftConfig.epochs
|
||
+ _sft_config_kwargs). Objective SATURATED by ~epoch 0.27 (loss~0, grad~0, acc 1.0, margins~27 flat — the
|
||
off-policy qwopus rejected pairs are trivially separable), so epochs ~0.3→3 were pure over-optimization + a
|
||
~6.5h outage. No mid-run checkpoint (save_steps 500 > total steps; save only at end) → killing lost the run.
|
||
FIX (3 edits to deployed harness, mtf-dev to canonicalize): `DpoConfig.epochs: float = 1` (mirrors SftConfig,
|
||
float for fractions); `_dpo_config_kwargs` now passes `"num_train_epochs": cfg.epochs`; recipe `dpo.epochs: 0.3`.
|
||
GPU-free verified (dpo.epochs=0.3 → num_train_epochs=0.3). Relaunched at 0.3 epoch (~30min precompute + ~12min
|
||
train = ~45min). **DONE + SANITY-CHECKED (exit 0, ~70min wall: ~30min fixed precompute + 45 steps @ ~51s/step;
|
||
train_loss 1.4e-5 @ epoch 0.301).** Fresh `data/spike/dpo_adapter/adapter_model.safetensors` (123MB) + checkpoint-45
|
||
banked. **3-way greedy sanity (base vs SFT vs DPO, via peft load + disable_adapter/set_adapter on GPU0):
|
||
ALL THREE DISTINCT** (base≠sft≠dpo) → full SFT→DPO pipeline applies end-to-end at inference. **DPO 0.3ep is
|
||
COHERENT, fluent, NOT degenerate** (early-stop avoided over-optimization) but the quality delta on a neutral
|
||
literary prompt is SUBTLE (DPO shares SFT's structure — it continues from it — with minor stylistic drift,
|
||
arguably slightly MORE genre-clichéd). Verdict: mechanics proven, quality gain modest as predicted for 0.3ep
|
||
on off-policy pairs; the real unlock remains on-policy rejected regen + on-domain (explicit E-RP) eval +
|
||
the LitBench/holdout run. gen+rp RESTORED healthy. Next: serve fork (SGLang-finish vs merge) on the DPO
|
||
adapter — same rekey_lora_for_vllm.py (zero-z) applies unchanged (mtf-dev confirmed).
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-08-t1-dpo-leg-is-running-unblocked-2-fixes
|
||
|
||
- `[2026-07-08]` **T1 DPO leg is RUNNING (unblocked) — 2 fixes applied to deployed backend.py.**
|
||
Blocker resolved: (1) **mtf-dev's v0.0.42 stub** `_stub_missing_optional_integrations` (last-resort sys.meta_path
|
||
finder → missing mergekit/llm_blender/weave resolve to MagicMock, never called → zero numerics risk; applied
|
||
VERBATIM to deployed `src/model_training_forge/train/backend.py` after `_unsloth_available()` + call-site before
|
||
`from trl import DPOTrainer`); (2) **my cosmetic `warnings_issued` shim** (trl-0.24 DPOTrainer.__init__:405 does
|
||
`model.warnings_issued["estimate_tokens"]=True` for warning-suppression; custom Qwen3_5 class under transformers
|
||
5.5.0 lacks the attr → `if not hasattr(model,"warnings_issued"): model.warnings_issued={}` before the
|
||
DPOTrainer(...).train() at backend.py:305 — cosmetic, zero training impact). Both edits are on the DEPLOYED
|
||
un-git'd copy only → **mtf-dev must canonicalize the warnings_issued shim into their repo** (told them). DPO
|
||
confirmed training: model loaded (851 shards), full 1196 pairs processed, in precompute_ref_log_probs (GPU0 93%
|
||
util, 54.8GB). Completion watcher armed (bg task) → restore gen+rp + verify dpo_adapter + ping mtf-dev on exit.
|
||
gen+rp STOPPED for the run (authorized window). Output → data/spike/dpo_adapter.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
### 2026-07-08-t1-dpo-leg-launch-prior-block-now-resolved
|
||
|
||
- `[2026-07-08]` **T1 DPO leg launch — prior BLOCK (now resolved above), kept for the launch recipe.**
|
||
Operator authorized the full DPO stage (via mtf-dev) + went AFK 2h. **PROVEN LAUNCH RECIPE** (replicates the
|
||
SFT container `aeon-t1-sft` exactly, only `--stage sft`→`dpo`): `sudo docker run -d --name aeon-t1-dpo
|
||
--entrypoint python3 --gpus all -e CUDA_VISIBLE_DEVICES=0 -e MTF_FORCE_TRL=1 -e PYTHONPATH=/mtf/src
|
||
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True -v /home/lkraven/model-training-forge:/mtf -v /tank:/tank
|
||
-w /mtf aeon-trainer:latest -u scripts/train.py --recipe recipes/training/qwen-3.5-122b-erp-lora/train.aeon-27b.yaml
|
||
--stage dpo`. **CRITICAL: `--entrypoint python3` is REQUIRED** — aeon-trainer's default entrypoint is
|
||
`["vllm","serve"]` (FROM vllm/vllm-openai) → without the override it runs vllm + hits a torch-ABI crash.
|
||
Dataset verified (pairs_dataset=train.flat.json=1196 pairs). **THE BLOCK:** `from trl import DPOTrainer`
|
||
(backend.py:256) eagerly pulls TRL 0.24.0's WHOLE optional-integration set — cascade: mergekit(missing)→
|
||
immutables→**mergekit-0.1.4↔pydantic-2.13 HARD incompat** (needs pydantic==2.10.6)→llm_blender→dataclasses_json→
|
||
**llm_blender-0.0.2↔transformers-5.5.0 HARD incompat** (TRANSFORMERS_CACHE removed, needs source patch)→weave→
|
||
(more). NONE used by our pair-based DPO. `pip install mergekit` w/deps is UNSAFE (downgrades accelerate
|
||
1.14→1.6). Safe partial recipe derived (core libs held: torch2.10/tf5.5.0/trl0.24.0/peft0.19.1/accel1.14.0)
|
||
but non-convergent → TRULY BLOCKING per operator's carve-out. Did NOT force-hack the proven training image.
|
||
Handed full diagnosis + recommended fix (lazy-import TRL patch, opt b) to mtf-dev (thread 01KWZG8GJX,
|
||
expects-reply, monitor armed). gen+rp RESTORED healthy. Relaunch = 1 min once mtf-dev delivers a working image.
|
||
|
||
_Archived 2026-08-03._
|
||
|
||
|
||
## Recent decisions (archived 2026-08-05 batch)
|
||
|
||
- `[2026-07-16]` **GPU re-org: char-rp→GPU1 + both cards re-optimized for max context.** Moved char-rp (Magidonia-24B) GPU0→GPU1, then maxed context: char-rp-reasoning 150K→256K (util 0.46, 1.56x), gen→256K + seqs 16→32 (util 0.42, 5.43x), granite 64K→**128K full-chapter** (util 0.27, 1.50x). FINAL: GPU0 ~14 G reserve (both seats 256K native), GPU1 ~6.7 G headroom. All healthy. LESSON: KV must hold ≥1× max-len (util-floor crashes) + per-model KV cost varies ~8× (MoE cheap, dense pricey) → tune util empirically.
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-16]` **granite right-sized → ~10.5 GB freed on GPU1** (util 0.34→0.18 + max-len 131072→65536; KV 6.45 GiB / 1.29x@65536). LESSON: ~950 MiB KV per 0.01 util here + KV must hold ≥1× max-len — util 0.15 crash-looped before 0.18 landed. `.env`-only, recreate `vllm-granite` alone (shared stack). [Superseded by the 07-16 GPU re-org above → 128K.]
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **image-bench eviction DONE (parked item closed).** Stopped vllm-qwen-image-bench (ana-ml2 GPU1, ~32 GB freed); LiteLLM `image-judge`+`qwen-image-bench` → gen :8015 (judge samplers + thinking-off); comfy-dev pinged; backfilled the canonical char-rp-reasoning litellm block. Revert ~90 s. auto-memory `project_arbo_gen_switch_imagebench_evict`.
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **Homepage AI-tab revamp** — flat "AI Systems" group → dedicated AI tab, 6 role-based groups + AI-Dormant; committed `569e1af`, pushed. (Also caught + pushed a ~100-commit unpushed eshpfi backlog.)
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **Home Assistant config repo created** (`vh/home-assistant-config`, private). UI-managed HA → allowlist model (YAML + curated secret-free `.storage` subset). git-in-place in `/config` on esh-docker-vm + scoped deploy key + local clone `~/development/home-assistant-config`.
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **char-rp-reasoning OOM rescue** — solo-restart on the packed GPU0 crash-looped; fixed via `expandable_segments:True` + util 0.39→0.38 + max-model-len 192K→150K. LESSON: `max-model-len` does NOT free vLLM VRAM (util-pinned KV pool). ~4.5 GB GPU0 headroom.
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **soong-lab `SOONG_LAB_LIBRARY_DIR` made persistent** (corviduo-dev) — was on the redeploy-wiped code default; set to `/home/infra-ops/soong-lab-data/library`, restarted. Closed a queued no-rush item.
|
||
_Archived 2026-08-05._
|
||
|
||
- `[2026-07-15]` **Statusline overhauled** (`~/.claude/statusline-command.sh`) — git state / 🔔🔕 monitor-armed / project tag / abs tokens / per-session cost / threshold-colored ctx+rate.
|
||
_Archived 2026-08-05._
|
||
|
||
## Tried and abandoned (archived) — moved 2026-08-12
|
||
|
||
- `[2026-07-01]` **MTP/spec-decode on a SHARED serving model helps single-stream but HURTS moderate-concurrency aggregate + silently ignores `min_p`/`logit_bias`** (qwopus `gen`: N=1 +12%, N=4 −20%). Reserve for dedicated/interactive deployments.
|
||
_Archived 2026-08-12._
|
||
|
||
- `[2026-07-02]` **irv-ml1 `/worktank` ROOT is root-owned — lkraven can't write there (irv-ml1 sudo needs a password) → stage model pulls to `/home`.** PIN THE A6000 BY UUID for training (native-CUDA ordering differs vs docker; the 3090 index 0 is usually near-full → OOM). `CUDA_VISIBLE_DEVICES=GPU-<uuid>`.
|
||
_Archived 2026-08-12._
|