memory: snapshot — 2026-06-19 gen model = Qwopus3.5-122B vision-intact NVFP4 LIVE on ana-ml2 GPU 0 (full 256K @ fp8 KV + CUDA graphs, util 0.95 + expandable_segments, 92.7 tok/s warm, 3.32x concurrency, text+image+video, tool-calling qwen3_coder; nightly+turboquant-4bit-KV proven UNNECESSARY — stable fp8 reaches 256K) replacing the bjk110 text-only qwen3.5-122b (which displaced mistral-small-4 → Worldtree character backend DARK until repointed, operator-acknowledged) + qwen-image-bench T2I judge replaced qwen3.6-35b-a3b on GPU 1 (alias image-judge) + TP=2 across both Blackwells REJECTED (PCIe-only PIX, no NVLink → all-reduce-bound, one-model-per-card is optimal; PP=2 only if a >96GB model is ever wanted) + foot-guns: MoE FusedMoE workspace is the ~3.1GB un-budgeted floor (can't fill to 0), discard cold tok/s reads (24.8 cold vs 92.7 warm).
This commit is contained in:
+27
-21
@@ -1,6 +1,6 @@
|
||||
# Persistent memory — eshpfi-management
|
||||
|
||||
_Last updated: 2026-06-18_
|
||||
_Last updated: 2026-06-19_
|
||||
|
||||
## Repo purpose
|
||||
|
||||
@@ -99,9 +99,9 @@ no longer deployed sidecars here. See Recent decisions.)
|
||||
|
||||
## Current state / in-flight
|
||||
|
||||
_As of 2026-06-18:_
|
||||
_As of 2026-06-19:_
|
||||
|
||||
- **Heretic abliterated Mistral Small 4 NVFP4 is LIVE as `mistral-small-4`** (ana-ml2 GPU 0) — executes the 06-16 "abliteration planned". Built in-house (darkc0de/…-heretic → NVFP4, vision tower kept bf16 → native format) and swapped in as the gateway `mistral-small-4` + `-reasoning` backend; A/B'd vs official, operator said **"heretic stays."** Official `mistral-small-4` stack staged-down (one-step revert: `down` heretic, `up` official). Build tooling `tools/mistral-small4-nvfp4/`. **PARKED (operator's call):** litellm comments + the official stack README still say "official NVFP4" (doc-drift); whether to notify char-role consumers the model is now abliterated. (auto-memory n/a; commits dd3a5c9/f566f61)
|
||||
- **Mistral Small 4 heretic is DISPLACED from GPU 0 (staged-down) — GPU 0 is now the `gen` model (Qwopus).** The heretic abliterated NVFP4 (`stacks/mistral-small-4-heretic/`, built in-house, "heretic stays", build tooling `tools/mistral-small4-nvfp4/`, byte-equiv to official) was downed 2026-06-19 to give GPU 0 to the gen swap. **CONSEQUENCE (operator-acknowledged, per the litellm ⚠️ note): the Worldtree demo+personal `character` backend — bound to `mistral-small-4` — is DARK until repointed.** To restore: down Qwopus + `up` the heretic stack (or serve it elsewhere). Prior doc-drift (official-stack README/comments saying "official NVFP4") now moot for GPU 0. (dd3a5c9, f566f61)
|
||||
|
||||
- **irv-ml1 VRAM consolidated — ComfyUI owns the full 48 GB A6000.** Pinned comfyui `NVIDIA_VISIBLE_DEVICES=1`; the audio/TTS zoo (chatterbox, parakeet live; vibevoice, yt-voice-clipper config-pinned; kokoro already there) moved to the 3090; downed dia2-2b (17-day stale), ace-step, csm-expressiva. comfy-dev torch-pin applied (`DISABLE_UPGRADES=true` @ torch 2.12.1, SageAttention rebuilt + matched). **WATCH:** the 3090 has ~18.7 GB free for the audio zoo — heavy *concurrent* on-demand audio could pressure it; vibevoice deploys from `/worktank/vibevoice/build` (pre-existing repo-vs-deploy drift). (a8550ad; auto-memory `reference_irv_ml1_comfyui_mmartial`)
|
||||
|
||||
@@ -117,24 +117,18 @@ _As of 2026-06-18:_
|
||||
|
||||
- **Worldtree demo + personal `character` model = mistral-small-4** (flipped from qwen3.6-35-a3b, 2026-06-16) — reordered `model_roles.yaml` `character.binds` mistral-first (first bind = default), qwen kept in the switch-allowlist; applied via pin-safe recreate, fresh-agent resolution verified. Backups `model_roles.yaml.bak-pre-mistral-character`.
|
||||
|
||||
- **ana-ml2 GPU layout RESHAPED again (2026-06-15/16) — both cards now full
|
||||
with NVFP4 tenants.** **GPU 0 = Mistral Small 4** (`mistral-small-4` stack,
|
||||
`mistralai/Mistral-Small-4-119B-2603-NVFP4`, 119B/6.5B-active MoE, :8010,
|
||||
gateway `mistral-small-4` + `mistral-small-4-reasoning`@effort=high). Pinned
|
||||
**vLLM v0.22.0** — the LAST release with working Mistral *vision* (#44911
|
||||
`fetch_images` regression breaks it on 0.22.1+/0.23.0). Serves the full native
|
||||
**256K context** (max-model-len 262144, max-num-seqs 32 — fits the tight card,
|
||||
~5 GB free). Text + vision both work; reasoning via `reasoning_effort` (BINARY:
|
||||
none|high). Dedicated single-tenant; the operator's creative-writing model
|
||||
(abliteration planned → it succeeds llama-swap). **GPU 1 = qwen36 swapped
|
||||
FP8→NVFP4** (`nvidia/Qwen3.6-35B-A3B-NVFP4`, fp16 KV, util 0.34, :8007, gateway
|
||||
name `qwen3.6-35b-a3b` UNCHANGED + `-thinking` variant) — the ModelOpt NVFP4
|
||||
MoE LOADS on 0.23.0 now (the 2026-06-14 "blocked" finding is RESOLVED). +
|
||||
**granite restored** (0.34/131072) + **Selene FP8 judge added** (`selene-1-mini-8b`,
|
||||
AtlaAI Selene-1-Mini-Llama-3.1-8B dynamic fp8, util 0.17, :8011, ctx 32768) +
|
||||
embed/rerank/reward. ~5.6 GB free. Prefix-caching ON on all 4 generative.
|
||||
**llama-swap is DOWN** (decommissioned from GPU 0 for Mistral; its qwen GGUF
|
||||
consumers migrated to the gateway). (auto-memory `reference_nvfp4_moe_loads_on_vllm_023`)
|
||||
- **ana-ml2 GPU layout (2026-06-19) — both 96GB Blackwells full, ONE model per
|
||||
card.** **GPU 0 = Qwopus3.5-122B-A10B** (the `gen`/`gen-reasoning` model;
|
||||
OpenYourMind Kimi-distilled abliterated NVFP4, VISION-INTACT MoE; `stacks/qwopus3.5-122b/`,
|
||||
:8013, served-name `qwen3.5-122-a10b`). Full **256K** (262144) @ fp8 KV + CUDA
|
||||
graphs, util 0.95 + `expandable_segments`, **92.7 tok/s** warm, 3.32x concurrency
|
||||
@256K, text+image+video, tool-calling `qwen3_coder`. **GPU 1 = qwen-image-bench**
|
||||
(T2I quality JUDGE, NVFP4; replaced qwen3.6-35b-a3b; alias `image-judge`, :8014) +
|
||||
granite-4.1-8b + selene-1-mini-8b + embed + rerank + reward — packed ~90.5/96 GB.
|
||||
**Interconnect = PCIe only (PIX, NO NVLink)** → one-model-per-card is the DELIBERATE
|
||||
optimal layout (zero cross-card traffic); TP=2 rejected this session (see Recent
|
||||
decisions). mistral-small-4 (heretic) + qwen3.6-35b-a3b both displaced.
|
||||
(20e796c, bfae924; auto-memory `reference_nvfp4_moe_loads_on_vllm_023`)
|
||||
|
||||
- **arbo engine builds handed to comfy-dev; Gitea Actions runner LIVE on irv-ml1.**
|
||||
Operator approved comfy-dev owning arbo engine deploys (`deploy-engine.sh`,
|
||||
@@ -222,6 +216,12 @@ _As of 2026-06-18:_
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-06-19]` **`gen` model → Qwopus3.5-122B-A10B (vision-intact NVFP4), full 256K @ fp8.** OpenYourMind/Qwopus3.5-122B-A10B-Kimi-K2.6-destilled-abliterated-NVFP4 on ana-ml2 GPU 0, served-name `qwen3.5-122-a10b` (gen/gen-reasoning/qwen-large route unchanged). Replaced the bjk110 text-only qwen3.5-122b (which had replaced mistral-small-4 earlier same-day). **KEY FINDING: the STABLE vLLM image + fp8 KV reaches the full 262144 — nightly+turboquant-4bit-KV was UNNECESSARY** (the hybrid SSM+attn KV pool is small; 11GB fp8 = 870k tokens = 3.32x concurrency @256K). graphs ON → 92.7 tok/s warm; util 0.95 + `expandable_segments` (0.96 OOMs the FusedMoE workspace); text+image+video + tool-calling qwen3_coder all verified. (20e796c, 5b06514)
|
||||
|
||||
- `[2026-06-19]` **TP=2 across the two ana-ml2 Blackwells REJECTED** (operator asked; recommended against). `nvidia-smi topo -m` = `PIX` (PCIe single-bridge, **NO NVLink** — datacenter-only). TP all-reduces ~twice/layer over PCIe (~64GB/s vs NVLink ~900GB/s) → all-reduce-bound → SLOWER for models that already fit + would evict the 6 GPU-1 services. **One-model-per-card is the optimal layout for non-NVLinked cards** (zero cross-card traffic). If a single >96GB model is ever wanted, the path is PIPELINE parallelism (PP=2, 1 hop/token) + GPU-1 relocation — NOT TP. (untracked by operator choice; "keep it there")
|
||||
|
||||
- `[2026-06-19]` **qwen-image-bench (T2I quality judge, NVFP4) replaced qwen3.6-35b-a3b on GPU 1** (operator), aliased `image-judge`; comfy-dev/arbo repointed off the killed `qwen3.6-35b-a3b` name. (bfae924, 5dfce04)
|
||||
|
||||
- `[2026-06-18]` **heretic abliterated Mistral Small 4 NVFP4 built + LIVE as `mistral-small-4`** (executes the 06-16 "abliteration planned"). `darkc0de/Mistral-Small-4-119B-2603-heretic` → in-house NVFP4 (vision bf16, `device_map=cpu`) → native format (HF Mistral4 is unserveable on vLLM) → drop-in stack `stacks/mistral-small-4-heretic/` under the same `--served-model-name mistral-small-4` (zero litellm change). A/B'd vs official (refusal+ability); operator: "heretic stays." Empirically byte-equivalent to the official NVFP4 (70.80 GB tensors, identical quant scope). (dd3a5c9, f566f61, `tools/mistral-small4-nvfp4/`)
|
||||
|
||||
- `[2026-06-18]` **irv-ml1 VRAM consolidation + comfy-dev torch-pin** (operator) — ComfyUI pinned to the A6000 exclusively (48 GB), audio zoo → 3090, downed dia2-2b/ace-step/csm-expressiva. comfy-dev's torch-pin: `DISABLE_UPGRADES=true` @ torch 2.12.1, SageAttention rebuilt against it. (a8550ad)
|
||||
@@ -317,6 +317,12 @@ _79 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-06-19]` **A MoE card can't be filled to 0 bytes free — the FusedMoE transient workspace is the floor.** vLLM's FusedMoE kernel allocates a ~3.09 GB transient workspace OUTSIDE its `gpu-memory-utilization` budget, into free VRAM, during graph capture + inference. util 0.96 OOM'd by 0.1 GB on it (`tried to allocate 3.09 GiB, 2.99 free`), worsened by 4.2 GB of PyTorch reserved-but-unallocated fragmentation. FIX: `PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True` (reclaims the fragmentation) + leave ~3.2 GB free (util ≤ ~0.95 on a tight-fit MoE). The workspace size is fixed; "fill to 0" is physically impossible for MoE.
|
||||
|
||||
- `[2026-06-19]` **vLLM decode tok/s: ALWAYS discard the first generation (cold warmup).** Cold single read = 24.8 tok/s; warm steady-state = 92.7 (3 runs identical). A cold read undersells decode ~3–4× — the first gen pays graph-replay/JIT warmup. Measure run 2+ over a ≥256-token output.
|
||||
|
||||
- `[2026-06-19]` **For full native 256K on one 96GB card, nightly+turboquant-4bit-KV was unnecessary for the Qwopus MoE.** Stable fp8 KV already fits 262144 — the hybrid SSM+attention model caches KV only on its attention layers, so the pool is small (11 GB fp8 = 870k tokens). 4-bit turboquant KV (nightly-only) only buys MORE concurrency, at a long-context-recall risk + FA2 fallback (incompatible with FA3). Reach for fp8 first; 4-bit only if you need heavy concurrency at long context.
|
||||
|
||||
- `[2026-06-18]` **mmartial `comfyui-nvidia-docker` image: root pip installs CRASH-LOOP the container.** `docker exec -u 0 pip install` (the documented node-install pattern) leaves root-owned files in the uid-1000 venv; the image's boot script re-manages that venv AS uid 1000 (its torch-upgrade step) → `Permission denied` on `setuptools/__pycache__` → `Torch installation failed` → crash loop (looks like a torch bug, is ownership). FIX: `chown -R 1000:1000 /comfy/mnt/venv` after any root install (host: `/worktank/comfyui/run/venv`; if crash-looping too fast to exec, `docker stop` → chown host path → `start`). The image ALSO auto-upgrades torch every boot (`USE_PIPUPGRADE`) → compiled exts drift; pin with `DISABLE_UPGRADES=true`. (auto-memory `reference_irv_ml1_comfyui_mmartial`)
|
||||
|
||||
- `[2026-06-17]` **Mistral HF→NVFP4 quant: the model-placement knob is the whole game.** `device_map="auto"` fills GPU0 → OOM during MoE un-fusing; constraining with `max_memory` offloads experts to the *meta* device → `Cannot copy out of meta tensor`. The working config is `device_map="cpu"` (CPU-resident model, sequential pipeline onloads each layer to GPU0). Plus: read shards with plain `read()` + `safetensors.torch.load(bytes)`, NOT `safe_open` (mmaps the whole shard → ENOMEM on `/tank` ZFS for the 50 GB shard, regardless of free RAM/overcommit). And llm-compressor's NVFP4 output KEEPS the `model.` prefix (not prefix-shifted).
|
||||
|
||||
Reference in New Issue
Block a user