From a67d4950d0250993dfcae2a42b1c9c04c955e7dc Mon Sep 17 00:00:00 2001 From: Vuong Hoang Date: Thu, 18 Jun 2026 13:48:12 -0700 Subject: [PATCH] =?UTF-8?q?memory:=20snapshot=20=E2=80=94=202026-06-18=20h?= =?UTF-8?q?eretic=20abliterated=20Mistral=20Small=204=20NVFP4=20BUILT=20+?= =?UTF-8?q?=20LIVE=20as=20mistral-small-4=20(in-house=20quant=20device=5Fm?= =?UTF-8?q?ap=3Dcpu=20=E2=86=92=20native-format=20convert=20=E2=86=92=20dr?= =?UTF-8?q?op-in=20stack=20under=20same=20served-name,=20A/B'd=20vs=20offi?= =?UTF-8?q?cial,=20operator=20"heretic=20stays";=20byte-equivalent=20to=20?= =?UTF-8?q?official=20NVFP4)=20+=20irv-ml1=20VRAM=20consolidation=20(Comfy?= =?UTF-8?q?UI=20pinned=20to=20A6000=20exclusive/48GB,=20audio=20zoo?= =?UTF-8?q?=E2=86=923090,=20downed=20dia/ace-step/csm,=20comfy-dev=20torch?= =?UTF-8?q?-pin=20DISABLE=5FUPGRADES@2.12.1=20+=20SageAttention=20rebuilt)?= =?UTF-8?q?=20+=20ComfyUI=209-node=20accel=20set=20installed=20for=20comfy?= =?UTF-8?q?-dev=20+=20ana-ml2=20durable=20vm.overcommit=5Fmemory=3D1=20+?= =?UTF-8?q?=20GLM5.2=20wired=20+=20nh3-extdev=20sudo-less=20manager=20box?= =?UTF-8?q?=20+=20/opt/externs=20pi-on-GLM=20client=20workspaces.?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Lessons: mmartial-comfyui root-install-leaves-root-owned-venv-files → boot-script crash-loop (chown -R 1000:1000 fix) + torch-upgrade-on-boot (DISABLE_UPGRADES); mistral HF→NVFP4 quant device_map=cpu (auto OOMs, constrained→meta-tensor) + non-mmap shard reads (safe_open mmap ENOMEMs on /tank ZFS) + NVFP4 keeps the model. prefix; HF-format Mistral4 UNSERVEABLE on vLLM (native mandatory); ComfyUI 0.24.1-not-0.19.3 version-drift kills module-level node imports + tensorrt-defaults-cu13-vs-cu12.9. Archived the [2026-06-14] cluster (11 entries: 6 decisions + 5 foot-guns; kept the still-active credential-migration directive, infra-ops litellm key, gitea-internal-route). --- archival-memory.md | 33 +++++++++++++++++++++++++ persistent-memory.md | 58 ++++++++++++++++++++++++-------------------- 2 files changed, 65 insertions(+), 26 deletions(-) diff --git a/archival-memory.md b/archival-memory.md index 07d8b77..2902e16 100644 --- a/archival-memory.md +++ b/archival-memory.md @@ -432,6 +432,24 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re - `[2026-06-09]` **LiteLLM scoped virtual keys issued to consumers** (operator-authorized): `brokkr-smithy` (all-proxy-models), `arbo-prompt-enhance` (comfy-dev — granite, later extended to qwen-vision). Mint via `/key/generate` (master `sk-corvid`), scope-restricted + rotatable, value → 600 file never the bus. (auto-memory `reference_litellm_gateway`) _Archived 2026-06-16._ +- `[2026-06-14]` **ana-ml2 GPU-1 vision upgraded: Qwen3.5-9B → Qwen3.6-35B-A3B (official FP8), served under its TRUE name only.** `qwen36-vl` replaces `qwen35-vl` on :8007 (`a0fed13`). The stale `qwen3.5-9b-fp8` name is KILLED at vLLM AND the litellm gateway (404/400) — a model is NEVER aliased under a prior model's name (silent substitution = downstream footgun; operator directive). Consumer comfy-dev/arbo migrated; arbo vkeys → all-proxy-models; shared `all-agents-local` key repointed qwen3.5-9b-fp8 → qwen3.6-35b-a3b. GPU-1 rebalanced for the ~34 GB FP8 weights (granite 0.35→0.24/64K; embed/rerank 0.05→0.03, reclaimed ~4 GB util-waste). Validated: vision correct, 20-concurrent = no OOM. (auto-memory `feedback_no_false_model_aliases`) + _Archived 2026-06-18._ + +- `[2026-06-14]` **NVFP4 was the lighter fit (~21 GB) but is BLOCKED on vLLM — FP8 is the working vision path.** `nvidia/Qwen3.6-35B-A3B-NVFP4` won't load: the ModelOpt-NVFP4-MoE loader errors on expert/lm_head scale keys across 0.19.1 (`w2_input_scale`) AND 0.22.0 (`lm_head.input_scale`, vllm #44081) — a pattern across modelopt NVFP4 MoEs. Revisit NVFP4 (frees ~13 GB on GPU 1) once fixed; the 21 GB checkpoint stays cached on ana-ml2. **(SUPERSEDED 2026-06-16 — it loads on vLLM 0.23.0; qwen36 swapped to NVFP4. See the top of this section.)** + _Archived 2026-06-18._ + +- `[2026-06-14]` **llama-swap qwen3.5-9b GPU-0 pin DROPPED; GPU 0 reserved for a creative-writing model (pick DEFERRED by operator).** Deep-research (this session) on big-fast-uncensored creative for a 96 GB Blackwell: **GLM-Steam-106B-A12B** (already in the llama-swap config — balanced default) vs **TheDrummer/Behemoth-X-123B-v2** (prose-tier, tops UGI writing+willingness) vs XORTRON-123B (max willingness, weak prose); GGUF-on-llama-swap is the serving path. Tracking: this session + llama-swap config (GLM-Steam present, `untracked by operator choice`). + _Archived 2026-06-18._ + +- `[2026-06-14]` **R16 splice-pivot yield probe executed** (infra-ops ran the irv-ml1 inference for brokkr; brokkr owns design + analysis). See Current state. Tracking: althing thread `01KV010WGS…`, `gen_yield_probe.py` in `irv-ml1:~/r16-vmoan-harness`. + _Archived 2026-06-18._ + +- `[2026-06-14]` **R16 vmoan inline-generation arc CLOSED — v1 at default decode (rep1.2/temp0.8) is the final Chatterbox-tag inline artifact.** Operator's ear rejected every alternative: v2/v3 windowing (omission vs coherence-loss), v4 multi-tag (cohesion held but lost to capacity-competition), emergent inline-token modulation (degenerates, not modulates), and the gen-time decode-polish sweep (soft tamers cut the NVV itself — same omission family as v2; p0 baseline beat p1). All adapters v1–v4 + `tokenizer.json.v3bak` preserved on `irv-ml1:~/r16-vmoan-harness`. Likely-next direction (deferred, NOT formalized): generate→bin→splice + one-shot-clone NVV pipeline routing around the inline-coherence wall. Tracking: brokkr R16 journal + althing thread `01KV010WGSSMPWRNCPAGSPK15Y`. + _Archived 2026-06-18._ + +- `[2026-06-14]` **Arbo deploy pipeline fixed, hardened, and version-controlled.** Prod rebuilt v0.11.1 → **v0.11.6** backend; the webhook machinery (`arbo-deploy.sh` + `arbo-webhook.py`, :9009 HMAC listener) is now repo-tracked at `stacks/arbo/` (was host-only = recoverability foot-gun). Deploy reaches gitea via the INTERNAL route (`10.250.50.70:222`) and restarts the engine ONLY on `catalog/` changes (graphs/frontend per-request; warn on `src/`|`Dockerfile` only — pyproject/uv.lock churn every commit). Operator kept arbo stack ownership in **eshpfi** (not migrated to comfy-dev's repo). Secret + `.env` stay host-only. Tracking: `6d66bc2`, `6e58e57`, `stacks/arbo/README` Q5. + _Archived 2026-06-18._ + ## Tried and abandoned (archived) - `[2026-04-30]` task-board workflow with @@ -874,3 +892,18 @@ _Entries moved out of persistent-memory.md to keep the active file scannable. Re - `[2026-06-11]` **A completion-poll `while pgrep -f ` SELF-MATCHES its own remote shell argv** — its own `pgrep -f` always finds itself → the loop never exits. Use a match pattern ABSENT from the poll command (the python stage, or a sentinel file), not the driver's own name. _Archived 2026-06-16._ + +- `[2026-06-14]` **vLLM ModelOpt-NVFP4-MoE loader is broken for current multimodal MoEs.** `nvidia/Qwen3.6-35B-A3B-NVFP4` fails weight-load: `KeyError: layers.0.mlp.experts.w2_input_scale` on 0.19.1, `lm_head.input_scale not registered` on 0.22.0 (vllm #44081); same class hits Gemma-4 MoE / Qwen3-30B-A3B NVFP4. The arch + quant ARE recognized (gets past arch resolution + vision-processor load) — it's the per-expert/lm_head scale-key mapping. Don't chase nightlies; use official FP8 until fixed. + _Archived 2026-06-18._ + +- `[2026-06-14]` **vLLM sampler-warmup OOMs on a shared GPU even when weights fit** — it warms the sampler with `max_num_seqs` (default **1024**) dummy requests, and a big vocab (Qwen3.6 = 248K) makes that a huge transient logits tensor. A vision endpoint doesn't need 1024-way concurrency: set `--max-num-seqs 32`. Separately, post-load `ValueError: No available memory for the cache blocks` means util is too thin (weights+activation+graph ate it) — for 34 GB FP8 weights, util ≥ ~0.45 to leave KV room. + _Archived 2026-06-18._ + +- `[2026-06-14]` **Recreating multiple vLLM services concurrently races the memory-profiling assertion** — `AssertionError: Error in memory profiling. Initial free memory X / current Y … other processes … release GPU memory while vLLM is profiling`. Recreate co-tenant vLLM services ONE AT A TIME (force-recreate one, wait healthy, next). + _Archived 2026-06-18._ + +- `[2026-06-14]` **embed/rerank (0.6B) at util 0.05 reserve ~5.5 GB each — mostly util-reservation WASTE, not need.** A 0.6B model needs ~1.2 GB weights + ~2.5 GB CUDA/torch context; util 0.03 (~3.6 GB) fits with room, reclaiming ~4 GB (vLLM reserves the util fraction regardless of actual KV; embedding models barely use KV). Real-need floor ~3 GB — don't go to 0.02. + _Archived 2026-06-18._ + +- `[2026-06-14]` **Chatterbox-Turbo decode-knob foot-guns** (R16 v1-polish + emergent probes): the turbo length cap is `max_gen_len` (default 1000) on `t3.inference_turbo`, NOT `max_new_tokens` — and `tts_turbo.generate` does NOT forward it (wrap inference_turbo to cap). `rep_pen 2.0 / temp 0.5` BACKFIRES (degenerate 24 s run-on). Soft decode tamers cut the NVV ITSELF, not just the run-on tail (operator: "p1 trims the moaning too") — gen-time polish can't beat v1's defaults. Inline base-NVV tokens DEGENERATE (moan-cascade + gibberish), they don't modulate the surrounding words. + _Archived 2026-06-18._ diff --git a/persistent-memory.md b/persistent-memory.md index 3b1598b..3cf4fa2 100644 --- a/persistent-memory.md +++ b/persistent-memory.md @@ -1,6 +1,6 @@ # Persistent memory — eshpfi-management -_Last updated: 2026-06-16_ +_Last updated: 2026-06-18_ ## Repo purpose @@ -99,7 +99,15 @@ no longer deployed sidecars here. See Recent decisions.) ## Current state / in-flight -_As of 2026-06-16:_ +_As of 2026-06-18:_ + +- **Heretic abliterated Mistral Small 4 NVFP4 is LIVE as `mistral-small-4`** (ana-ml2 GPU 0) — executes the 06-16 "abliteration planned". Built in-house (darkc0de/…-heretic → NVFP4, vision tower kept bf16 → native format) and swapped in as the gateway `mistral-small-4` + `-reasoning` backend; A/B'd vs official, operator said **"heretic stays."** Official `mistral-small-4` stack staged-down (one-step revert: `down` heretic, `up` official). Build tooling `tools/mistral-small4-nvfp4/`. **PARKED (operator's call):** litellm comments + the official stack README still say "official NVFP4" (doc-drift); whether to notify char-role consumers the model is now abliterated. (auto-memory n/a; commits dd3a5c9/f566f61) + +- **irv-ml1 VRAM consolidated — ComfyUI owns the full 48 GB A6000.** Pinned comfyui `NVIDIA_VISIBLE_DEVICES=1`; the audio/TTS zoo (chatterbox, parakeet live; vibevoice, yt-voice-clipper config-pinned; kokoro already there) moved to the 3090; downed dia2-2b (17-day stale), ace-step, csm-expressiva. comfy-dev torch-pin applied (`DISABLE_UPGRADES=true` @ torch 2.12.1, SageAttention rebuilt + matched). **WATCH:** the 3090 has ~18.7 GB free for the audio zoo — heavy *concurrent* on-demand audio could pressure it; vibevoice deploys from `/worktank/vibevoice/build` (pre-existing repo-vs-deploy drift). (a8550ad; auto-memory `reference_irv_ml1_comfyui_mmartial`) + +- **ComfyUI acceleration set (9 nodes) installed for comfy-dev on irv-ml1** — TeaCache, WaveSpeed(+FBCache), SageAttention2, Detail-Daemon(+bleh), PAG, dynamic-thresholding, Skimmed_CFG, TensorRT(cu12), SUPIR; comfy-dev wires + benchmarks. The full-A6000 re-tiered Flux-TRT + SUPIR from Ada-dormant to wire-today. + +- **nh3-extdev (10.100.50.42) = sudo-less infra-ops manager box** (successor to nh3-ansible) + `/opt/externs/{gbcnc,surefire,svsconstruction}` client workspaces for **pi-on-GLM-5.2** client agents. **GLM 5.2 wired into litellm** (`glm-5.2` + `-reasoning` via z.ai). - **LitBench-RM stood up → now ON-DEMAND / DOWN.** `SAA-Lab/Llama8B-CreativeWritingVerifier` (R19 creative-quality reward judge for brokkr/Dvalin) served on irv-ml1 A6000 via vLLM `--runner pooling` (`http://10.100.79.3:8202/classify`, raw passage text → scalar). brokkr-validated (does NOT penalize explicit content). Taken DOWN to on-demand (operator) — resident it held ~19.6 GB crowding comfyui's A6000 slot; weights staged on irv-ml1, ~90s respin (command in auto-memory). (auto-memory `reference_litbench_rm_irv_ml1`) @@ -214,6 +222,18 @@ _As of 2026-06-16:_ ## Recent decisions +- `[2026-06-18]` **heretic abliterated Mistral Small 4 NVFP4 built + LIVE as `mistral-small-4`** (executes the 06-16 "abliteration planned"). `darkc0de/Mistral-Small-4-119B-2603-heretic` → in-house NVFP4 (vision bf16, `device_map=cpu`) → native format (HF Mistral4 is unserveable on vLLM) → drop-in stack `stacks/mistral-small-4-heretic/` under the same `--served-model-name mistral-small-4` (zero litellm change). A/B'd vs official (refusal+ability); operator: "heretic stays." Empirically byte-equivalent to the official NVFP4 (70.80 GB tensors, identical quant scope). (dd3a5c9, f566f61, `tools/mistral-small4-nvfp4/`) + +- `[2026-06-18]` **irv-ml1 VRAM consolidation + comfy-dev torch-pin** (operator) — ComfyUI pinned to the A6000 exclusively (48 GB), audio zoo → 3090, downed dia2-2b/ace-step/csm-expressiva. comfy-dev's torch-pin: `DISABLE_UPGRADES=true` @ torch 2.12.1, SageAttention rebuilt against it. (a8550ad) + +- `[2026-06-18]` **ComfyUI acceleration set (9 nodes) installed for comfy-dev** on irv-ml1's `comfyui` (arbo's box) — work order delivered; comfy-dev wires + benchmarks (TeaCache→Wan first). + +- `[2026-06-17]` **ana-ml2 `vm.overcommit_memory=1` made durable** (sysctl drop-in, `playbooks/ana-ml2-overcommit-memory.yaml`) — overcommit=0 + zero swap caps CommitLimit at ~RAM/2; the resident vLLM services ate the headroom so a large model-file mmap ENOMEM'd despite ~393 GB free. (fc88eff) + +- `[2026-06-17]` **GLM 5.2 wired into litellm** (`glm-5.2` + `glm-5.2-reasoning`, z.ai passthrough, `extra_body.thinking.type` toggle — mirrors the GLM-5.1 split). (fe77a35) + +- `[2026-06-17]` **nh3-extdev stood up as a sudo-LESS infra-ops manager box** (operator) — key-only, password-locked, no NOPASSWD/docker (deliberately tighter than the fleet infra-ops identity); successor to nh3-ansible. Hosts `/opt/externs` client workspaces for pi-on-GLM-5.2 client agents. (a841eab; auto-memory `reference_infra_ops_sudo_identity`) + - `[2026-06-16]` **litellm `strip_empty_tools` pre-call hook shipped** (`d1bea13`) — an empty `tools:[]` 500s vLLM ("tools must not be an empty array"); a global `litellm_settings.callbacks` CustomLogger pops it (+ orphaned `tool_choice`) before forwarding, so it covers EVERY vLLM model, not one. `drop_params` only drops unsupported PARAMS, not empty VALUES. Mounts beside config.yaml (litellm resolves callbacks relative to the config dir). Verified live across granite/mistral/stream. (`stacks/litellm/conf/strip_empty_tools.py`) - `[2026-06-16]` **single-file `gateway-chat.html` playground shipped** (`984ca3d`, `tools/`) — zero-dep browser chat straight to the gateway (`:4000`, CORS open), system-prompt box, streaming SSE, renders `reasoning_content`, NEVER sends `tools`. Built because the LiteLLM admin-UI playground can't test vLLM-backed models (see Tried-and-abandoned). Serve on-request via `python3 -m http.server -d tools`. @@ -291,24 +311,20 @@ _As of 2026-06-16:_ - `[2026-06-14]` **LiteLLM infra-ops admin key provisioned** (operator) — resolves the LiteLLM half of the credential-migration directive; use it for ALL gateway ops (NOT `sk-corvid`). Value at `~/.config/litellm/infra-ops-key` (mode 600); gateway reachable directly from nh3-dev at 10.250.50.70:4000. (auto-memory `reference_litellm_infra_ops_key`) -- `[2026-06-14]` **ana-ml2 GPU-1 vision upgraded: Qwen3.5-9B → Qwen3.6-35B-A3B (official FP8), served under its TRUE name only.** `qwen36-vl` replaces `qwen35-vl` on :8007 (`a0fed13`). The stale `qwen3.5-9b-fp8` name is KILLED at vLLM AND the litellm gateway (404/400) — a model is NEVER aliased under a prior model's name (silent substitution = downstream footgun; operator directive). Consumer comfy-dev/arbo migrated; arbo vkeys → all-proxy-models; shared `all-agents-local` key repointed qwen3.5-9b-fp8 → qwen3.6-35b-a3b. GPU-1 rebalanced for the ~34 GB FP8 weights (granite 0.35→0.24/64K; embed/rerank 0.05→0.03, reclaimed ~4 GB util-waste). Validated: vision correct, 20-concurrent = no OOM. (auto-memory `feedback_no_false_model_aliases`) - -- `[2026-06-14]` **NVFP4 was the lighter fit (~21 GB) but is BLOCKED on vLLM — FP8 is the working vision path.** `nvidia/Qwen3.6-35B-A3B-NVFP4` won't load: the ModelOpt-NVFP4-MoE loader errors on expert/lm_head scale keys across 0.19.1 (`w2_input_scale`) AND 0.22.0 (`lm_head.input_scale`, vllm #44081) — a pattern across modelopt NVFP4 MoEs. Revisit NVFP4 (frees ~13 GB on GPU 1) once fixed; the 21 GB checkpoint stays cached on ana-ml2. **(SUPERSEDED 2026-06-16 — it loads on vLLM 0.23.0; qwen36 swapped to NVFP4. See the top of this section.)** - -- `[2026-06-14]` **llama-swap qwen3.5-9b GPU-0 pin DROPPED; GPU 0 reserved for a creative-writing model (pick DEFERRED by operator).** Deep-research (this session) on big-fast-uncensored creative for a 96 GB Blackwell: **GLM-Steam-106B-A12B** (already in the llama-swap config — balanced default) vs **TheDrummer/Behemoth-X-123B-v2** (prose-tier, tops UGI writing+willingness) vs XORTRON-123B (max willingness, weak prose); GGUF-on-llama-swap is the serving path. Tracking: this session + llama-swap config (GLM-Steam present, `untracked by operator choice`). - -- `[2026-06-14]` **R16 splice-pivot yield probe executed** (infra-ops ran the irv-ml1 inference for brokkr; brokkr owns design + analysis). See Current state. Tracking: althing thread `01KV010WGS…`, `gen_yield_probe.py` in `irv-ml1:~/r16-vmoan-harness`. - - `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.** Agents currently reuse the operator's PERSONAL creds for infra ops — vh Gitea **admin** via `tea` (used this session to mint a `read:package` token for ratatoskr, id 13, **under vh**), `sk-corvid` litellm master for vkey admin. Stand up service accounts (a `claude-bot` Gitea user + scoped tokens, a distinct litellm admin key); re-mint consumer creds under them; flag personal-cred fallbacks until done. (auto-memory `project_migrate_infra_access_to_claude_credentials`) -- `[2026-06-14]` **R16 vmoan inline-generation arc CLOSED — v1 at default decode (rep1.2/temp0.8) is the final Chatterbox-tag inline artifact.** Operator's ear rejected every alternative: v2/v3 windowing (omission vs coherence-loss), v4 multi-tag (cohesion held but lost to capacity-competition), emergent inline-token modulation (degenerates, not modulates), and the gen-time decode-polish sweep (soft tamers cut the NVV itself — same omission family as v2; p0 baseline beat p1). All adapters v1–v4 + `tokenizer.json.v3bak` preserved on `irv-ml1:~/r16-vmoan-harness`. Likely-next direction (deferred, NOT formalized): generate→bin→splice + one-shot-clone NVV pipeline routing around the inline-coherence wall. Tracking: brokkr R16 journal + althing thread `01KV010WGSSMPWRNCPAGSPK15Y`. - -- `[2026-06-14]` **Arbo deploy pipeline fixed, hardened, and version-controlled.** Prod rebuilt v0.11.1 → **v0.11.6** backend; the webhook machinery (`arbo-deploy.sh` + `arbo-webhook.py`, :9009 HMAC listener) is now repo-tracked at `stacks/arbo/` (was host-only = recoverability foot-gun). Deploy reaches gitea via the INTERNAL route (`10.250.50.70:222`) and restarts the engine ONLY on `catalog/` changes (graphs/frontend per-request; warn on `src/`|`Dockerfile` only — pyproject/uv.lock churn every commit). Operator kept arbo stack ownership in **eshpfi** (not migrated to comfy-dev's repo). Secret + `.env` stay host-only. Tracking: `6d66bc2`, `6e58e57`, `stacks/arbo/README` Q5. - -_73 older entries archived to archival-memory.md._ +_79 older entries archived to archival-memory.md._ ## Tried and abandoned +- `[2026-06-18]` **mmartial `comfyui-nvidia-docker` image: root pip installs CRASH-LOOP the container.** `docker exec -u 0 pip install` (the documented node-install pattern) leaves root-owned files in the uid-1000 venv; the image's boot script re-manages that venv AS uid 1000 (its torch-upgrade step) → `Permission denied` on `setuptools/__pycache__` → `Torch installation failed` → crash loop (looks like a torch bug, is ownership). FIX: `chown -R 1000:1000 /comfy/mnt/venv` after any root install (host: `/worktank/comfyui/run/venv`; if crash-looping too fast to exec, `docker stop` → chown host path → `start`). The image ALSO auto-upgrades torch every boot (`USE_PIPUPGRADE`) → compiled exts drift; pin with `DISABLE_UPGRADES=true`. (auto-memory `reference_irv_ml1_comfyui_mmartial`) + +- `[2026-06-17]` **Mistral HF→NVFP4 quant: the model-placement knob is the whole game.** `device_map="auto"` fills GPU0 → OOM during MoE un-fusing; constraining with `max_memory` offloads experts to the *meta* device → `Cannot copy out of meta tensor`. The working config is `device_map="cpu"` (CPU-resident model, sequential pipeline onloads each layer to GPU0). Plus: read shards with plain `read()` + `safetensors.torch.load(bytes)`, NOT `safe_open` (mmaps the whole shard → ENOMEM on `/tank` ZFS for the 50 GB shard, regardless of free RAM/overcommit). And llm-compressor's NVFP4 output KEEPS the `model.` prefix (not prefix-shifted). + +- `[2026-06-17]` **HF-format Mistral Small 4 is UNSERVEABLE on vLLM** — there is no HF `Mistral4` backbone in any vLLM version; it serves ONLY via the native loader (`--config-format/--load-format/--tokenizer-mode mistral` against params.json + consolidated*.safetensors + tekken.json). So a HF-format quant MUST be converted to native before it can serve. + +- `[2026-06-18]` **ComfyUI custom nodes break on version-assumption drift** — the box runs 0.24.1 (not the 0.19.3 a work order assumed); 0.24.1 refactored `precompute_freqs_cis` into class methods, and TeaCache imports it at MODULE level → kills the WHOLE node (guard the LTX-only import). Also `pip install tensorrt` defaults to **cu13** libs against a cu12.9 stack → use `tensorrt-cu12`. + - `[2026-06-16]` **litellm 500 `Router.acompletion()/aembedding() missing 'messages'/'input'` = a request missing `Content-Type: application/json`, NOT a gateway outage.** curl `-d` defaults to form-encoding → litellm can't parse the JSON body → `data` reaches the router without `messages`/`input` → 500 (should be a 400; litellm #16993). My own diagnostic calls dropped the header → I misread it as a gateway outage and needlessly bounced the gateway ~4× chasing a phantom (image/version/config were fine throughout; a malformed UI-added "Mistral Story Eval" model in the DB was a red herring I deleted). ALWAYS send `-H "Content-Type: application/json"` testing litellm; reproduce with a header'd call before declaring a litellm incident. - `[2026-06-16]` **LiteLLM admin-UI playground can't test vLLM-backed models** — it auto-sends empty `tools:[]`, vLLM 400s (litellm #6228); the gateway `strip_empty_tools` hook is a PROXY hook and structurally can't reach the UI's in-process `litellm.completion()` call. Off-ramp = `tools/gateway-chat.html`. (Langfuse playground also out: its SSRF guard blocks internal-IP LLM connections, wontfix Langfuse #13097.) (auto-memory `reference_litellm_ui_playground_vllm_deadend`) @@ -358,16 +374,6 @@ _73 older entries archived to archival-memory.md._ - `[2026-06-15]` **`pkill -f althing-light-monitor` SELF-MATCHES the killing shell** (the pattern is in the command's own argv) → kills itself mid-run (exit 144/truncated output). Stop the light-monitor via `althing-cli stop-monitor` or a captured PID — never `pkill -f `. The singleton lock can also RACE to 2 live monitors during re-arm churn; keep exactly one tracked (run_in_background) monitor, and a raw `&` monitor is untracked (no harness fire-notification — don't use it). -- `[2026-06-14]` **vLLM ModelOpt-NVFP4-MoE loader is broken for current multimodal MoEs.** `nvidia/Qwen3.6-35B-A3B-NVFP4` fails weight-load: `KeyError: layers.0.mlp.experts.w2_input_scale` on 0.19.1, `lm_head.input_scale not registered` on 0.22.0 (vllm #44081); same class hits Gemma-4 MoE / Qwen3-30B-A3B NVFP4. The arch + quant ARE recognized (gets past arch resolution + vision-processor load) — it's the per-expert/lm_head scale-key mapping. Don't chase nightlies; use official FP8 until fixed. - -- `[2026-06-14]` **vLLM sampler-warmup OOMs on a shared GPU even when weights fit** — it warms the sampler with `max_num_seqs` (default **1024**) dummy requests, and a big vocab (Qwen3.6 = 248K) makes that a huge transient logits tensor. A vision endpoint doesn't need 1024-way concurrency: set `--max-num-seqs 32`. Separately, post-load `ValueError: No available memory for the cache blocks` means util is too thin (weights+activation+graph ate it) — for 34 GB FP8 weights, util ≥ ~0.45 to leave KV room. - -- `[2026-06-14]` **Recreating multiple vLLM services concurrently races the memory-profiling assertion** — `AssertionError: Error in memory profiling. Initial free memory X / current Y … other processes … release GPU memory while vLLM is profiling`. Recreate co-tenant vLLM services ONE AT A TIME (force-recreate one, wait healthy, next). - -- `[2026-06-14]` **embed/rerank (0.6B) at util 0.05 reserve ~5.5 GB each — mostly util-reservation WASTE, not need.** A 0.6B model needs ~1.2 GB weights + ~2.5 GB CUDA/torch context; util 0.03 (~3.6 GB) fits with room, reclaiming ~4 GB (vLLM reserves the util fraction regardless of actual KV; embedding models barely use KV). Real-need floor ~3 GB — don't go to 0.02. - - `[2026-06-14]` **Fleet/colo hosts must reach gitea over the INTERNAL route, NOT the public IP.** `gitea.phasefinal.com` = public `38.120.12.44` (ana-srv1); gitea is a container on ana-docker, git-SSH `10.250.50.70:222` + HTTP `:3000`. A fleet host egressing to public `:22` gets fail2ban-banned after any retrying git loop → silently wedges webhook auto-deploys (`git fetch` times out under `set -euo pipefail`, aborts before reset). Bit irv-ml1's arbo deploy. `:22` on `10.250.50.70` is ana-docker's HOST sshd (deploy key → Permission denied), NOT gitea. Documented `docs/orientation.md` (`6e58e57`). -- `[2026-06-14]` **Chatterbox-Turbo decode-knob foot-guns** (R16 v1-polish + emergent probes): the turbo length cap is `max_gen_len` (default 1000) on `t3.inference_turbo`, NOT `max_new_tokens` — and `tts_turbo.generate` does NOT forward it (wrap inference_turbo to cap). `rep_pen 2.0 / temp 0.5` BACKFIRES (degenerate 24 s run-on). Soft decode tamers cut the NVV ITSELF, not just the run-on tail (operator: "p1 trims the moaning too") — gen-time polish can't beat v1's defaults. Inline base-NVV tokens DEGENERATE (moan-cascade + gibberish), they don't modulate the surrounding words. - -_65 older entries archived to archival-memory.md._ +_70 older entries archived to archival-memory.md._