Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md
T
vh 4390be947d feat(flash-next-seat): serve orcarouter weight-only NVFP4 on gen-large
Swaps gen-large from the dealignai ModelOpt W4A4 build to
orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4, which is weight-only on both
axes (W8 float attn, W4 float experts, input_activations: null) and so avoids
the 4-bit-activation long-context degradation mode.

The checkpoint was previously recorded as unloadable on any mainline vLLM,
requiring a from-source PLE-loader patch. That conclusion was wrong on cost.
Qwen4ExpPLEEmbeddingMethod.from_quant_config checks ple_embedding_dtype as
branch 1, before any quant-config type check, and its NotImplementedError for
CompressedTensorsConfig is scoped to the PLE path only -- experts and dense
load through the ordinary compressed-tensors paths. Verified by instantiating
the real config and calling the selector both ways before doing any work.

orcarouter ships a bf16 PLE, so the fix was to make the declaration true:
convert the 51.2B-param table to FP8 and declare it. Its 128 PLE tensors sit
in one shard file with nothing else in it. Global amax 0.0894, per-shard
outlier ratio 1.66x, scale chosen exactly representable in bf16 so no
scale-rounding error stacks on quantization; amax maps to 446.17/448, no
clipping. Round-trip 2.655% RMS relative, 0.002% underflow, 0 saturation --
the same FP8-PLE treatment dealignai already shipped. MTP head (31 tensors)
and vision tower carried through untouched.

A second, independent blocker followed: orcarouter labels its 12 QSA layers
qwen_sparse_attention, which vLLM rejects; it accepts full_attention and
selects QSA via indexer_n_heads. Confirmed indexer_n_heads == 4 in both this
and the dealignai checkpoint before renaming -- without that check the rename
silently selects plain attention and serves a subtly wrong model that still
passes a healthcheck.

Measured on the live seat: healthy, coherent, KV 344,155 tokens @ 262,144 ctx,
MTP k=3 at 60.4% acceptance / 2.81 mean acceptance length, warm decode median
167.5 tok/s at conc=1 (n=5, spread 12.2%). The reorg note's dealignai figure
came from a different harness, so this is not claimed as a win over it; what
it does establish is that weight-only experts did not cost decode speed.

Still open: controlled quality A/B vs dealignai, and a deep-prefill probe at
262K. Rollback is two .env keys; dealignai remains on disk.

Also corrects the README's MTP-is-off section, stale since k=3 was deployed,
and adds a superseded-claims row to the quantization playbook.
2026-09-14 02:48:42 -07:00

7.6 KiB
Raw Blame History

2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker

An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything here is on the running box; regenerate the authoritative view with scripts/seat-inventory.py (reads the live containers). Cross-refs: 2026-09-13-flash-next-seat-and-fv-outage.

Seat topology now (2026-09-14 ~01:40 PT)

GPU seats
0 mog-sec (sec :8019, dflash k=7) · sentinel-r3 (sentinel-r3 :8025, dflash k=7 — NEW)
1 meromero-charrp (char-rp :8016, MeroMero-v2-31B dense — restored) · erp-seat (char-rp-fast :8021) · reward · coder · embed · rerank
2 flash-next (gen-large :8022) — UP on orcarouter (PLE converted bf16→FP8; MTP k=3, 60.4% accept)
3 RESERVED scratch — empty, operator directive; benches/quants/probes only

What changed tonight

  1. flash-next gained MTP k=3. Campaign in services/flash-next-mtp-bench/ measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34% k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user → conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights; 14 GiB OOMs). Warm decode ~121 tok/s. stacks/flash-next-seat/compose.yaml.
  2. gen consolidation. All 8 gen/gen-reasoning/summarizer/summarizer-large/ classifier/chat-judge/image-judge/qwen-image-bench LiteLLM entries repointed to flash-next (:8022); the 27B dense gen seat RETIRED, freeing 38.4 GB on GPU0. ⚠ judge aliases now score against different weights — prior scores incomparable.
  3. char-rp restored to MeroMero-v2-31B. Was serving a leftover-test RedHatAI 26B MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back. char-rp-fast (:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so char-rp stays the quality seat. stacks/meromero-charrp/.
  4. Sentinel-R3 A/B + dflash. glyphsoftware/sentinel-r3 (proprietary license — operator's call) is a REAL SFT pentest finetune vs mog-sec's persona-on-stock. Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's finetuned body: dflash 2.40 vs MTP 2.18 mean acceptance length (+11%), warm decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over. stacks/sentinel-r3/. ⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot + concurrent-contention artifact; warm+isolated it was 121. Operator caught it by testing the running sec (102 tok/s) as reference.

✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed)

gen-large now serves orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 from /tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8. Healthy, coherent, MTP k=3.

The earlier diagnosis was right about the symptom and wrong about the cost. It said the only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell build. Reading the loader in the running nightly showed otherwise:

from_quant_config  (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168)
  1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method   <-- BEFORE any type check
  2. quant_config is None                   -> unquantized
  3. ModelOptMixedPrecisionConfig           -> FP8 / unquantized
  4. ModelOptQuantConfigBase + excluded     -> unquantized
  5. not isinstance(quant_config, Fp8Config)-> NotImplementedError   <-- the blocker

Branch 1 is unconditional, and the NotImplementedError is scoped to the PLE embedding path only — experts and dense layers of a compressed-tensors qwen4_exp build load through vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built the real CompressedTensorsConfig from orca's own config and called from_quant_config both ways — as-shipped raises, with the declaration returns Qwen4ExpPLEFp8EmbeddingMethod.

What was actually done

  1. Converted the PLE table bf16 → FP8. orca's 128 PLE tensors live in exactly ONE shard (model-00002-of-00017.safetensors, 95.4 GiB) with no other tensors in it — a clean split. Converted on GPU3 (reserved scratch) in 139 s into 8 model-plefp8-* files.
    • global amax 0.0894; per-shard outlier ratio only 1.66x, so one global scale fits
    • scale chosen exactly representable in bf16 (2.002716e-04) so no scale-rounding error stacks on the quantization error; amax maps to 446.17 / 448 → no clipping
    • round-trip 2.655 % RMS relative, 0.002 % underflow, 0 saturation
    • weight_scale written BF16 [1], matching gorbatjovy's published format (read from its actual safetensors header, not guessed)
    • MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched
  2. Declared it: text_config.ple_embedding_dtype = "float8_e4m3fn".
  3. Renamed layer_types: orca labels its 12 QSA layers qwen_sparse_attention; vLLM accepts only linear_attention / full_attention and picks QSA via indexer_n_heads. ⚠ Verified indexer_n_heads == 4 in BOTH orca and dealignai before renaming — without it the rename silently selects PLAIN attention and serves a subtly wrong model that still looks healthy. Every QSA/indexer key matches dealignai exactly.
  4. .env: FN_MODEL → the converted dir, FN_QUANT → compressed-tensors.

Measured on the live seat

warm decode, conc=1, greedy 300 tok, n=5 median 167.5 tok/s (min 150.0, max 170.4, spread 12.2 %)
MTP k=3 acceptance 60.4 %, mean acceptance length 2.81 (per-pos 80.6/60.8/40.8 %)
KV 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB
on card ~75 GiB; PLE 47.7 GiB pinned host RAM

⚠ Do NOT read 167.5 as a win over dealignai. dealignai's ~121 tok/s in this file came from a different harness/prompt; cross-harness comparison is invalid. What IS established is that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did not cost decode speed, which was the standing risk of giving up FP4 tensor-core compute.

Still open on this seat

  • Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context degradation axis). Needs a controlled harness + noise floor.
  • Deep-prefill probe at 262K against THIS checkpoint. Startup is not a depth test.
  • Rollback is two .env keys; .env.bak-preorca-20260914-023408 on the host.
  • Disk: the convert copied ~75 GiB of unchanged shards because hardlinks hit EXDEV (separate bind mounts of the same fs). Harmless; reclaimable by relinking if /tank tightens.

Other open items

  • cyberprev quant (3rd sec candidate hotdogs/Qwen3.8-27B-abliterated-cyber-preview, bf16 at /tank/aimodels/cyberprev-bf16): restart crashed on a transformers head-count config error in quant-work/.venv — the SAME venv that reached 49/65 on Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path (per services/gen-seat-mixed-quant/README.md), not the venv. Retry that way.
  • sec vs sentinel-r3 quality A/B — both live and gateway-callable; operator to judge.
  • ⚠ Suspect vault entry: fv-gateway/infra-ops-password is 16 chars matching a boot-UUID prefix exactly — possibly malformed. Eyeball.
  • os-nut not installed on the FV OPNsense — with the firewall now on the 5P1000 UPS, a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today.
  • Branch breaker rating + 4-card ammeter reading still open — every power table is arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.