Swaps gen-large from the dealignai ModelOpt W4A4 build to orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4, which is weight-only on both axes (W8 float attn, W4 float experts, input_activations: null) and so avoids the 4-bit-activation long-context degradation mode. The checkpoint was previously recorded as unloadable on any mainline vLLM, requiring a from-source PLE-loader patch. That conclusion was wrong on cost. Qwen4ExpPLEEmbeddingMethod.from_quant_config checks ple_embedding_dtype as branch 1, before any quant-config type check, and its NotImplementedError for CompressedTensorsConfig is scoped to the PLE path only -- experts and dense load through the ordinary compressed-tensors paths. Verified by instantiating the real config and calling the selector both ways before doing any work. orcarouter ships a bf16 PLE, so the fix was to make the declaration true: convert the 51.2B-param table to FP8 and declare it. Its 128 PLE tensors sit in one shard file with nothing else in it. Global amax 0.0894, per-shard outlier ratio 1.66x, scale chosen exactly representable in bf16 so no scale-rounding error stacks on quantization; amax maps to 446.17/448, no clipping. Round-trip 2.655% RMS relative, 0.002% underflow, 0 saturation -- the same FP8-PLE treatment dealignai already shipped. MTP head (31 tensors) and vision tower carried through untouched. A second, independent blocker followed: orcarouter labels its 12 QSA layers qwen_sparse_attention, which vLLM rejects; it accepts full_attention and selects QSA via indexer_n_heads. Confirmed indexer_n_heads == 4 in both this and the dealignai checkpoint before renaming -- without that check the rename silently selects plain attention and serves a subtly wrong model that still passes a healthcheck. Measured on the live seat: healthy, coherent, KV 344,155 tokens @ 262,144 ctx, MTP k=3 at 60.4% acceptance / 2.81 mean acceptance length, warm decode median 167.5 tok/s at conc=1 (n=5, spread 12.2%). The reorg note's dealignai figure came from a different harness, so this is not claimed as a win over it; what it does establish is that weight-only experts did not cost decode speed. Still open: controlled quality A/B vs dealignai, and a deep-prefill probe at 262K. Rollback is two .env keys; dealignai remains on disk. Also corrects the README's MTP-is-off section, stale since k=3 was deployed, and adds a superseded-claims row to the quantization playbook.
7.6 KiB
2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker
An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything
here is on the running box; regenerate the authoritative view with
scripts/seat-inventory.py (reads the live containers). Cross-refs:
2026-09-13-flash-next-seat-and-fv-outage.
Seat topology now (2026-09-14 ~01:40 PT)
| GPU | seats |
|---|---|
| 0 | mog-sec (sec :8019, dflash k=7) · sentinel-r3 (sentinel-r3 :8025, dflash k=7 — NEW) |
| 1 | meromero-charrp (char-rp :8016, MeroMero-v2-31B dense — restored) · erp-seat (char-rp-fast :8021) · reward · coder · embed · rerank |
| 2 | flash-next (gen-large :8022) — UP on orcarouter (PLE converted bf16→FP8; MTP k=3, 60.4% accept) |
| 3 | RESERVED scratch — empty, operator directive; benches/quants/probes only |
What changed tonight
- flash-next gained MTP k=3. Campaign in
services/flash-next-mtp-bench/measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34% k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user → conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights; 14 GiB OOMs). Warm decode ~121 tok/s.stacks/flash-next-seat/compose.yaml. - gen consolidation. All 8
gen/gen-reasoning/summarizer/summarizer-large/classifier/chat-judge/image-judge/qwen-image-benchLiteLLM entries repointed to flash-next (:8022); the 27B dense gen seat RETIRED, freeing 38.4 GB on GPU0. ⚠ judge aliases now score against different weights — prior scores incomparable. - char-rp restored to MeroMero-v2-31B. Was serving a leftover-test RedHatAI 26B
MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back.
char-rp-fast(:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so char-rp stays the quality seat.stacks/meromero-charrp/. - Sentinel-R3 A/B + dflash.
glyphsoftware/sentinel-r3(proprietary license — operator's call) is a REAL SFT pentest finetune vs mog-sec's persona-on-stock. Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's finetuned body: dflash 2.40 vs MTP 2.18 mean acceptance length (+11%), warm decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over.stacks/sentinel-r3/. ⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot + concurrent-contention artifact; warm+isolated it was 121. Operator caught it by testing the runningsec(102 tok/s) as reference.
✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed)
gen-large now serves orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 from
/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8. Healthy, coherent, MTP k=3.
The earlier diagnosis was right about the symptom and wrong about the cost. It said the only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell build. Reading the loader in the running nightly showed otherwise:
from_quant_config (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168)
1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method <-- BEFORE any type check
2. quant_config is None -> unquantized
3. ModelOptMixedPrecisionConfig -> FP8 / unquantized
4. ModelOptQuantConfigBase + excluded -> unquantized
5. not isinstance(quant_config, Fp8Config)-> NotImplementedError <-- the blocker
Branch 1 is unconditional, and the NotImplementedError is scoped to the PLE embedding
path only — experts and dense layers of a compressed-tensors qwen4_exp build load through
vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built
the real CompressedTensorsConfig from orca's own config and called from_quant_config
both ways — as-shipped raises, with the declaration returns Qwen4ExpPLEFp8EmbeddingMethod.
What was actually done
- Converted the PLE table bf16 → FP8. orca's 128 PLE tensors live in exactly ONE shard
(
model-00002-of-00017.safetensors, 95.4 GiB) with no other tensors in it — a clean split. Converted on GPU3 (reserved scratch) in 139 s into 8model-plefp8-*files.- global amax 0.0894; per-shard outlier ratio only 1.66x, so one global scale fits
- scale chosen exactly representable in bf16 (2.002716e-04) so no scale-rounding error stacks on the quantization error; amax maps to 446.17 / 448 → no clipping
- round-trip 2.655 % RMS relative, 0.002 % underflow, 0 saturation
weight_scalewritten BF16 [1], matching gorbatjovy's published format (read from its actual safetensors header, not guessed)- MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched
- Declared it:
text_config.ple_embedding_dtype = "float8_e4m3fn". - Renamed
layer_types: orca labels its 12 QSA layersqwen_sparse_attention; vLLM accepts onlylinear_attention/full_attentionand picks QSA viaindexer_n_heads. ⚠ Verifiedindexer_n_heads == 4in BOTH orca and dealignai before renaming — without it the rename silently selects PLAIN attention and serves a subtly wrong model that still looks healthy. Every QSA/indexer key matches dealignai exactly. .env:FN_MODEL→ the converted dir,FN_QUANT→compressed-tensors.
Measured on the live seat
| warm decode, conc=1, greedy 300 tok, n=5 | median 167.5 tok/s (min 150.0, max 170.4, spread 12.2 %) |
| MTP k=3 | acceptance 60.4 %, mean acceptance length 2.81 (per-pos 80.6/60.8/40.8 %) |
| KV | 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB |
| on card | ~75 GiB; PLE 47.7 GiB pinned host RAM |
⚠ Do NOT read 167.5 as a win over dealignai. dealignai's ~121 tok/s in this file came from a different harness/prompt; cross-harness comparison is invalid. What IS established is that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did not cost decode speed, which was the standing risk of giving up FP4 tensor-core compute.
Still open on this seat
- Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context degradation axis). Needs a controlled harness + noise floor.
- Deep-prefill probe at 262K against THIS checkpoint. Startup is not a depth test.
- Rollback is two
.envkeys;.env.bak-preorca-20260914-023408on the host. - Disk: the convert copied ~75 GiB of unchanged shards because hardlinks hit
EXDEV(separate bind mounts of the same fs). Harmless; reclaimable by relinking if /tank tightens.
Other open items
- cyberprev quant (3rd sec candidate
hotdogs/Qwen3.8-27B-abliterated-cyber-preview, bf16 at/tank/aimodels/cyberprev-bf16): restart crashed on a transformers head-count config error inquant-work/.venv— the SAME venv that reached 49/65 on Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path (perservices/gen-seat-mixed-quant/README.md), not the venv. Retry that way. - sec vs sentinel-r3 quality A/B — both live and gateway-callable; operator to judge.
- ⚠ Suspect vault entry:
fv-gateway/infra-ops-passwordis 16 chars matching a boot-UUID prefix exactly — possibly malformed. Eyeball. - os-nut not installed on the FV OPNsense — with the firewall now on the 5P1000 UPS, a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today.
- Branch breaker rating + 4-card ammeter reading still open — every power table is arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.