# 2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything here is on the running box; regenerate the authoritative view with `scripts/seat-inventory.py` (reads the live containers). Cross-refs: [[2026-09-13-flash-next-seat-and-fv-outage]]. ## Seat topology now (2026-09-14 ~01:40 PT) | GPU | seats | |---|---| | 0 | `mog-sec` (`sec` :8019, dflash k=7) · **`sentinel-r3`** (`sentinel-r3` :8025, dflash k=7 — NEW) | | 1 | **`meromero-charrp`** (`char-rp` :8016, MeroMero-v2-31B dense — restored) · `erp-seat` (`char-rp-fast` :8021) · reward · coder · embed · rerank | | 2 | `flash-next` (`gen-large` :8022) — **UP on orcarouter** (PLE converted bf16→FP8; MTP k=3, 60.4% accept) | | 3 | **RESERVED scratch** — empty, operator directive; benches/quants/probes only | ## What changed tonight 1. **flash-next gained MTP k=3.** Campaign in `services/flash-next-mtp-bench/` measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34% k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user → conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights; 14 GiB OOMs). Warm decode ~121 tok/s. `stacks/flash-next-seat/compose.yaml`. 2. **gen consolidation.** All 8 `gen`/`gen-reasoning`/`summarizer`/`summarizer-large`/ `classifier`/`chat-judge`/`image-judge`/`qwen-image-bench` LiteLLM entries repointed to flash-next (:8022); the **27B dense gen seat RETIRED**, freeing 38.4 GB on GPU0. ⚠ judge aliases now score against different weights — prior scores incomparable. 3. **char-rp restored to MeroMero-v2-31B.** Was serving a leftover-test RedHatAI 26B MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back. `char-rp-fast` (:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so char-rp stays the quality seat. `stacks/meromero-charrp/`. 4. **Sentinel-R3 A/B + dflash.** `glyphsoftware/sentinel-r3` (proprietary license — operator's call) is a REAL SFT pentest finetune vs mog-sec (M.O.G.-SEC), which is ALSO a finetune — a refusal-free offense+defense cyber SFT, NOT a persona-on-stock as earlier notes claimed. All three sec seats are Qwen3.8-27B finetunes; they differ in focus. Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's finetuned body: **dflash 2.40 vs MTP 2.18 mean acceptance length (+11%)**, warm decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over. `stacks/sentinel-r3/`. ⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot + concurrent-contention artifact; warm+isolated it was 121. Operator caught it by testing the running `sec` (102 tok/s) as reference. ## ✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed) gen-large now serves **`orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4`** from `/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8`. Healthy, coherent, MTP k=3. **The earlier diagnosis was right about the symptom and wrong about the cost.** It said the only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell build. Reading the loader in the running nightly showed otherwise: ``` from_quant_config (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168) 1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method <-- BEFORE any type check 2. quant_config is None -> unquantized 3. ModelOptMixedPrecisionConfig -> FP8 / unquantized 4. ModelOptQuantConfigBase + excluded -> unquantized 5. not isinstance(quant_config, Fp8Config)-> NotImplementedError <-- the blocker ``` Branch 1 is unconditional, and the `NotImplementedError` is **scoped to the PLE embedding path only** — experts and dense layers of a compressed-tensors qwen4_exp build load through vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built the real `CompressedTensorsConfig` from orca's own config and called `from_quant_config` both ways — as-shipped raises, with the declaration returns `Qwen4ExpPLEFp8EmbeddingMethod`. ### What was actually done 1. **Converted the PLE table bf16 → FP8.** orca's 128 PLE tensors live in exactly ONE shard (`model-00002-of-00017.safetensors`, 95.4 GiB) with **no other tensors in it** — a clean split. Converted on GPU3 (reserved scratch) in 139 s into 8 `model-plefp8-*` files. - global amax **0.0894**; per-shard outlier ratio only **1.66x**, so one global scale fits - scale chosen **exactly representable in bf16** (2.002716e-04) so no scale-rounding error stacks on the quantization error; amax maps to **446.17 / 448** → no clipping - round-trip **2.655 % RMS relative**, 0.002 % underflow, **0 saturation** - `weight_scale` written **BF16 [1]**, matching gorbatjovy's published format (read from its actual safetensors header, not guessed) - MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched 2. **Declared it**: `text_config.ple_embedding_dtype = "float8_e4m3fn"`. 3. **Renamed `layer_types`**: orca labels its 12 QSA layers `qwen_sparse_attention`; vLLM accepts only `linear_attention` / `full_attention` and picks QSA via `indexer_n_heads`. ⚠ **Verified `indexer_n_heads == 4` in BOTH orca and dealignai before renaming** — without it the rename silently selects PLAIN attention and serves a subtly wrong model that still looks healthy. Every QSA/indexer key matches dealignai exactly. 4. `.env`: `FN_MODEL` → the converted dir, `FN_QUANT` → `compressed-tensors`. ### Measured on the live seat | | | |---|---| | warm decode, conc=1, greedy 300 tok, **n=5** | median **167.5 tok/s** (min 150.0, max 170.4, spread 12.2 %) | | MTP k=3 | acceptance **60.4 %**, mean acceptance length **2.81** (per-pos 80.6/60.8/40.8 %) | | KV | 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB | | on card | ~75 GiB; PLE 47.7 GiB pinned host RAM | ⚠ **Do NOT read 167.5 as a win over dealignai.** dealignai's ~121 tok/s in this file came from a different harness/prompt; cross-harness comparison is invalid. ⚠ The 167.5 may itself have been contended — the operator was using the seat around that window — so treat it as a LOWER BOUND, not a clean solo figure. Operator's own session reported **140 tok/s average** in real use while the 258K depth probe was running against the same card: an independent, contended floor that agrees with the picture. What IS established is that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did **not** cost decode speed, which was the standing risk of giving up FP4 tensor-core compute. ### Still open on this seat - Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context degradation axis). Needs a controlled harness + noise floor. - ✅ **Deep-prefill probe DONE 2026-09-14: clean to 258,517 tokens.** Non-repeating prompts, 6 depths 32K->258K, all 200 OK, and **zero allocator OOM/CUBLAS/illegal-memory in the engine log** — the detector that caught dealignai's 155K near-miss. Run under real operator load, so a stricter test than solo. Positive control passed (a ~265K prompt got a clean 400 naming the limit). #54919 did not reproduce: 258K prefilled in 28.9 s, ~8,900 tok/s, near-linear. ⚠ The probe's MEMORY column was blind and must not be reused: `--kv-cache-memory` pins the pool ("skipped memory profiling"), so GPU use is flat vs depth, and before/after `nvidia-smi` bracketing cannot see a transient mid-prefill spike. Identical readings across an 8x depth range were the tell. Peak-activation headroom remains UNMEASURED. ⭐ Calibration for re-runs: random hex words tokenize at **7.9 tok/word**. - ⚠⚠ **NO LOCAL ROLLBACK.** dealignai weights DELETED 2026-09-14 on operator instruction (125 GiB reclaimed). `.env.bak-preorca-20260914-023408` still names the old paths but they no longer exist — it is a record, not a revert. Reverting = 126 GiB re-download. The pristine 170 GiB `qwen38-flash-next-orcarouter-nvfp4` IS retained (redo the conversion from it; do not delete it without a reason). - ✅ **Disk reclaimed 2026-09-14: 76 GiB.** The 28 non-PLE shards duplicated between the pristine and converted orca dirs are now **hardlinked** (294G apparent → 218G actual). All 28 verified **byte-identical by SHA-256** first, then `ln` to a temp name + atomic `rename` over the target — never rm-then-ln, which leaves a window with no file. Done live with the seat serving; it never blinked. ⚠ **The two dirs now SHARE INODES.** Editing a shared file *in place* in either dir changes both. Shards are never edited in place, and `config.json` / `model.safetensors.index.json` are deliberately NOT shared (the conversion changed them) — but a future session must copy-then-edit, not edit in place. ## Other open items - **cyberprev quant** (3rd sec candidate `hotdogs/Qwen3.8-27B-abliterated-cyber-preview`, bf16 at `/tank/aimodels/cyberprev-bf16`): restart crashed on a transformers head-count config error in `quant-work/.venv` — the SAME venv that reached 49/65 on Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path (per `services/gen-seat-mixed-quant/README.md`), not the venv. Retry that way. - **sec vs sentinel-r3 quality A/B** — both live and gateway-callable; operator to judge. - ⚠ **Suspect vault entry**: `fv-gateway/infra-ops-password` is 16 chars matching a boot-UUID prefix exactly — possibly malformed. Eyeball. - **os-nut not installed** on the FV OPNsense — with the firewall now on the 5P1000 UPS, a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today. - **Branch breaker rating + 4-card ammeter reading** still open — every power table is arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.