memory: snapshot — 2026-06-14 FP8 vision cutover (qwen35-vl→qwen36-vl, truthful naming, GPU-1 rebalance) + llama-swap pin drop + R16 yield probe executed + standing credential-migration directive; NVFP4-on-vLLM-blocked + sampler-warmup/profiling-race/embed-rerank-waste lessons
This commit is contained in:
+35
-10
@@ -113,17 +113,24 @@ _As of 2026-06-14:_
|
||||
the internal-route repoint made it MOOT; ban-cleanup is optional hygiene
|
||||
(status unverified).
|
||||
|
||||
- **R16 vmoan pilot is CLOSED** — v1 (default decode) is the final
|
||||
Chatterbox-tag inline artifact (see Recent decisions). The only forward
|
||||
thread: a splice-pivot **de-risk yield-probe** brokkr surfaced 2026-06-14
|
||||
(45 standalone-NVV clips + speaker-embedding/acoustic metrics) —
|
||||
**surfaced to operator, awaiting go**; it needs Resemblyzer/ECAPA +
|
||||
librosa/parselmouth installs in the harness venv.
|
||||
- **R16 splice-pivot yield probe RAN 2026-06-14** (was "awaiting go"). Inline-gen
|
||||
arc stays CLOSED (v1 @ default decode); the splice pivot's cheap de-risk is now
|
||||
executed: 45 standalone NVV ([moan_soft]/[moan_intense]/[groan] × seeds 1-15)
|
||||
from Eleanor.wav on the v4 adapter + objective metrics (Resemblyzer spk-cosine,
|
||||
MFCC-dist, F0, parselmouth HNR) → `/mnt/smithy/scratch/r16_audition_yield`,
|
||||
served on brokkr-audition :8137. 45/45, 0 degenerate; spk-cosine 0.48-0.81.
|
||||
Brokkr notified. NEXT: operator ear-bin (usable/impure) → brokkr's yield% /
|
||||
survivorship / identity analysis. `gen_yield_probe.py` in `irv-ml1:~/r16-vmoan-harness`.
|
||||
|
||||
- **ana-ml2 on Blackwell** (dual RTX PRO 6000 Blackwell Max-Q, 96 GB ea,
|
||||
cc 12.0). GPU 0 held free for large-model hot-loads (llama-swap pinned);
|
||||
GPU 1 steady tenants — granite 131k + qwen3.5-vl 65k + embed/rerank/reward
|
||||
trio, ~3.5 GB free. CLAUDE.md GPU-spec doc-fix LANDED (`355a240`).
|
||||
- **ana-ml2 GPU layout RESHAPED by the FP8 vision cutover** (2026-06-14, `a0fed13`;
|
||||
dual RTX PRO 6000 Blackwell Max-Q, 96 GB ea, cc 12.0). GPU 1 now: **Qwen3.6-35B-
|
||||
A3B-FP8 vision** (`qwen36-vl`, :8007, util 0.46 — replaced the Qwen3.5-9B) +
|
||||
granite (0.24 / 64K) + reward (0.10) + embed/rerank (0.03 ea) ≈ 90 GB, ~7.5 GB
|
||||
headroom (the OOM buffer — verified under 20-concurrent/endpoint load). **GPU 0
|
||||
now FULLY FREE** (llama-swap qwen3.5-9b pin dropped) — reserved for a big-fast-
|
||||
uncensored creative-writing model (pick DEFERRED, see Recent decisions). Vision
|
||||
served under its TRUE name only (`qwen3.6-35b-a3b`); the old `qwen3.5-9b-fp8`
|
||||
name is killed at vLLM + the gateway.
|
||||
|
||||
- **R17 v2 corpus characterization still on the local irv-ml1 branch, push
|
||||
HELD** (`r17-v2-characterization`). No-push rule AND ASR-content-exposure
|
||||
@@ -161,6 +168,16 @@ _As of 2026-06-14:_
|
||||
|
||||
## Recent decisions
|
||||
|
||||
- `[2026-06-14]` **ana-ml2 GPU-1 vision upgraded: Qwen3.5-9B → Qwen3.6-35B-A3B (official FP8), served under its TRUE name only.** `qwen36-vl` replaces `qwen35-vl` on :8007 (`a0fed13`). The stale `qwen3.5-9b-fp8` name is KILLED at vLLM AND the litellm gateway (404/400) — a model is NEVER aliased under a prior model's name (silent substitution = downstream footgun; operator directive). Consumer comfy-dev/arbo migrated; arbo vkeys → all-proxy-models; shared `all-agents-local` key repointed qwen3.5-9b-fp8 → qwen3.6-35b-a3b. GPU-1 rebalanced for the ~34 GB FP8 weights (granite 0.35→0.24/64K; embed/rerank 0.05→0.03, reclaimed ~4 GB util-waste). Validated: vision correct, 20-concurrent = no OOM. (auto-memory `feedback_no_false_model_aliases`)
|
||||
|
||||
- `[2026-06-14]` **NVFP4 was the lighter fit (~21 GB) but is BLOCKED on vLLM — FP8 is the working vision path.** `nvidia/Qwen3.6-35B-A3B-NVFP4` won't load: the ModelOpt-NVFP4-MoE loader errors on expert/lm_head scale keys across 0.19.1 (`w2_input_scale`) AND 0.22.0 (`lm_head.input_scale`, vllm #44081) — a pattern across modelopt NVFP4 MoEs. Revisit NVFP4 (frees ~13 GB on GPU 1) once fixed; the 21 GB checkpoint stays cached on ana-ml2.
|
||||
|
||||
- `[2026-06-14]` **llama-swap qwen3.5-9b GPU-0 pin DROPPED; GPU 0 reserved for a creative-writing model (pick DEFERRED by operator).** Deep-research (this session) on big-fast-uncensored creative for a 96 GB Blackwell: **GLM-Steam-106B-A12B** (already in the llama-swap config — balanced default) vs **TheDrummer/Behemoth-X-123B-v2** (prose-tier, tops UGI writing+willingness) vs XORTRON-123B (max willingness, weak prose); GGUF-on-llama-swap is the serving path. Tracking: this session + llama-swap config (GLM-Steam present, `untracked by operator choice`).
|
||||
|
||||
- `[2026-06-14]` **R16 splice-pivot yield probe executed** (infra-ops ran the irv-ml1 inference for brokkr; brokkr owns design + analysis). See Current state. Tracking: althing thread `01KV010WGS…`, `gen_yield_probe.py` in `irv-ml1:~/r16-vmoan-harness`.
|
||||
|
||||
- `[2026-06-14]` **STANDING DIRECTIVE: migrate ALL infra access to Claude-specific credentials.** Agents currently reuse the operator's PERSONAL creds for infra ops — vh Gitea **admin** via `tea` (used this session to mint a `read:package` token for ratatoskr, id 13, **under vh**), `sk-corvid` litellm master for vkey admin. Stand up service accounts (a `claude-bot` Gitea user + scoped tokens, a distinct litellm admin key); re-mint consumer creds under them; flag personal-cred fallbacks until done. (auto-memory `project_migrate_infra_access_to_claude_credentials`)
|
||||
|
||||
- `[2026-06-14]` **R16 vmoan inline-generation arc CLOSED — v1 at default decode (rep1.2/temp0.8) is the final Chatterbox-tag inline artifact.** Operator's ear rejected every alternative: v2/v3 windowing (omission vs coherence-loss), v4 multi-tag (cohesion held but lost to capacity-competition), emergent inline-token modulation (degenerates, not modulates), and the gen-time decode-polish sweep (soft tamers cut the NVV itself — same omission family as v2; p0 baseline beat p1). All adapters v1–v4 + `tokenizer.json.v3bak` preserved on `irv-ml1:~/r16-vmoan-harness`. Likely-next direction (deferred, NOT formalized): generate→bin→splice + one-shot-clone NVV pipeline routing around the inline-coherence wall. Tracking: brokkr R16 journal + althing thread `01KV010WGSSMPWRNCPAGSPK15Y`.
|
||||
|
||||
- `[2026-06-14]` **Arbo deploy pipeline fixed, hardened, and version-controlled.** Prod rebuilt v0.11.1 → **v0.11.6** backend; the webhook machinery (`arbo-deploy.sh` + `arbo-webhook.py`, :9009 HMAC listener) is now repo-tracked at `stacks/arbo/` (was host-only = recoverability foot-gun). Deploy reaches gitea via the INTERNAL route (`10.250.50.70:222`) and restarts the engine ONLY on `catalog/` changes (graphs/frontend per-request; warn on `src/`|`Dockerfile` only — pyproject/uv.lock churn every commit). Operator kept arbo stack ownership in **eshpfi** (not migrated to comfy-dev's repo). Secret + `.env` stay host-only. Tracking: `6d66bc2`, `6e58e57`, `stacks/arbo/README` Q5.
|
||||
@@ -215,6 +232,14 @@ _50 older entries archived to archival-memory.md._
|
||||
|
||||
## Tried and abandoned
|
||||
|
||||
- `[2026-06-14]` **vLLM ModelOpt-NVFP4-MoE loader is broken for current multimodal MoEs.** `nvidia/Qwen3.6-35B-A3B-NVFP4` fails weight-load: `KeyError: layers.0.mlp.experts.w2_input_scale` on 0.19.1, `lm_head.input_scale not registered` on 0.22.0 (vllm #44081); same class hits Gemma-4 MoE / Qwen3-30B-A3B NVFP4. The arch + quant ARE recognized (gets past arch resolution + vision-processor load) — it's the per-expert/lm_head scale-key mapping. Don't chase nightlies; use official FP8 until fixed.
|
||||
|
||||
- `[2026-06-14]` **vLLM sampler-warmup OOMs on a shared GPU even when weights fit** — it warms the sampler with `max_num_seqs` (default **1024**) dummy requests, and a big vocab (Qwen3.6 = 248K) makes that a huge transient logits tensor. A vision endpoint doesn't need 1024-way concurrency: set `--max-num-seqs 32`. Separately, post-load `ValueError: No available memory for the cache blocks` means util is too thin (weights+activation+graph ate it) — for 34 GB FP8 weights, util ≥ ~0.45 to leave KV room.
|
||||
|
||||
- `[2026-06-14]` **Recreating multiple vLLM services concurrently races the memory-profiling assertion** — `AssertionError: Error in memory profiling. Initial free memory X / current Y … other processes … release GPU memory while vLLM is profiling`. Recreate co-tenant vLLM services ONE AT A TIME (force-recreate one, wait healthy, next).
|
||||
|
||||
- `[2026-06-14]` **embed/rerank (0.6B) at util 0.05 reserve ~5.5 GB each — mostly util-reservation WASTE, not need.** A 0.6B model needs ~1.2 GB weights + ~2.5 GB CUDA/torch context; util 0.03 (~3.6 GB) fits with room, reclaiming ~4 GB (vLLM reserves the util fraction regardless of actual KV; embedding models barely use KV). Real-need floor ~3 GB — don't go to 0.02.
|
||||
|
||||
- `[2026-06-14]` **Fleet/colo hosts must reach gitea over the INTERNAL route, NOT the public IP.** `gitea.phasefinal.com` = public `38.120.12.44` (ana-srv1); gitea is a container on ana-docker, git-SSH `10.250.50.70:222` + HTTP `:3000`. A fleet host egressing to public `:22` gets fail2ban-banned after any retrying git loop → silently wedges webhook auto-deploys (`git fetch` times out under `set -euo pipefail`, aborts before reset). Bit irv-ml1's arbo deploy. `:22` on `10.250.50.70` is ana-docker's HOST sshd (deploy key → Permission denied), NOT gitea. Documented `docs/orientation.md` (`6e58e57`).
|
||||
|
||||
- `[2026-06-14]` **Chatterbox-Turbo decode-knob foot-guns** (R16 v1-polish + emergent probes): the turbo length cap is `max_gen_len` (default 1000) on `t3.inference_turbo`, NOT `max_new_tokens` — and `tts_turbo.generate` does NOT forward it (wrap inference_turbo to cap). `rep_pen 2.0 / temp 0.5` BACKFIRES (degenerate 24 s run-on). Soft decode tamers cut the NVV ITSELF, not just the run-on tail (operator: "p1 trims the moaning too") — gen-time polish can't beat v1's defaults. Inline base-NVV tokens DEGENERATE (moan-cascade + gibberish), they don't modulate the surrounding words.
|
||||
|
||||
Reference in New Issue
Block a user