Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md
T
vh 3906c6842c docs: correct sec-seat lineage — M.O.G.-SEC/mog-sec is an offense+defense SFT finetune, not a persona-on-stock
The sentinel-r3 header and two memory notes described mog-sec (Blackfrost
M.O.G.-SEC / Qwentium) as 'a persona system prompt on stock weights'. Its card
is explicit that it is NOT: base_model_relation: finetune on Qwen/Qwen3.8-27B,
a refusal-free offense+defense cybersecurity SFT with YaRN 1M context ('not a
system-prompt sticker on a stock Qwen'). So all three sec-seat candidates are
Qwen3.8-27B SFT finetunes and differ in training focus, not in kind:
mog-sec = broad offense+defense SFT; sentinel-r3 = pentest agent-trajectory SFT;
cyberprev = cyber tool-calling LoRA SFT on an abliterated base.
2026-09-14 07:31:23 -07:00

9.8 KiB
Raw Blame History

2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker

An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything here is on the running box; regenerate the authoritative view with scripts/seat-inventory.py (reads the live containers). Cross-refs: 2026-09-13-flash-next-seat-and-fv-outage.

Seat topology now (2026-09-14 ~01:40 PT)

GPU seats
0 mog-sec (sec :8019, dflash k=7) · sentinel-r3 (sentinel-r3 :8025, dflash k=7 — NEW)
1 meromero-charrp (char-rp :8016, MeroMero-v2-31B dense — restored) · erp-seat (char-rp-fast :8021) · reward · coder · embed · rerank
2 flash-next (gen-large :8022) — UP on orcarouter (PLE converted bf16→FP8; MTP k=3, 60.4% accept)
3 RESERVED scratch — empty, operator directive; benches/quants/probes only

What changed tonight

  1. flash-next gained MTP k=3. Campaign in services/flash-next-mtp-bench/ measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34% k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user → conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights; 14 GiB OOMs). Warm decode ~121 tok/s. stacks/flash-next-seat/compose.yaml.
  2. gen consolidation. All 8 gen/gen-reasoning/summarizer/summarizer-large/ classifier/chat-judge/image-judge/qwen-image-bench LiteLLM entries repointed to flash-next (:8022); the 27B dense gen seat RETIRED, freeing 38.4 GB on GPU0. ⚠ judge aliases now score against different weights — prior scores incomparable.
  3. char-rp restored to MeroMero-v2-31B. Was serving a leftover-test RedHatAI 26B MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back. char-rp-fast (:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so char-rp stays the quality seat. stacks/meromero-charrp/.
  4. Sentinel-R3 A/B + dflash. glyphsoftware/sentinel-r3 (proprietary license — operator's call) is a REAL SFT pentest finetune vs mog-sec (M.O.G.-SEC), which is ALSO a finetune — a refusal-free offense+defense cyber SFT, NOT a persona-on-stock as earlier notes claimed. All three sec seats are Qwen3.8-27B finetunes; they differ in focus. Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's finetuned body: dflash 2.40 vs MTP 2.18 mean acceptance length (+11%), warm decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over. stacks/sentinel-r3/. ⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot + concurrent-contention artifact; warm+isolated it was 121. Operator caught it by testing the running sec (102 tok/s) as reference.

✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed)

gen-large now serves orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4 from /tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8. Healthy, coherent, MTP k=3.

The earlier diagnosis was right about the symptom and wrong about the cost. It said the only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell build. Reading the loader in the running nightly showed otherwise:

from_quant_config  (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168)
  1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method   <-- BEFORE any type check
  2. quant_config is None                   -> unquantized
  3. ModelOptMixedPrecisionConfig           -> FP8 / unquantized
  4. ModelOptQuantConfigBase + excluded     -> unquantized
  5. not isinstance(quant_config, Fp8Config)-> NotImplementedError   <-- the blocker

Branch 1 is unconditional, and the NotImplementedError is scoped to the PLE embedding path only — experts and dense layers of a compressed-tensors qwen4_exp build load through vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built the real CompressedTensorsConfig from orca's own config and called from_quant_config both ways — as-shipped raises, with the declaration returns Qwen4ExpPLEFp8EmbeddingMethod.

What was actually done

  1. Converted the PLE table bf16 → FP8. orca's 128 PLE tensors live in exactly ONE shard (model-00002-of-00017.safetensors, 95.4 GiB) with no other tensors in it — a clean split. Converted on GPU3 (reserved scratch) in 139 s into 8 model-plefp8-* files.
    • global amax 0.0894; per-shard outlier ratio only 1.66x, so one global scale fits
    • scale chosen exactly representable in bf16 (2.002716e-04) so no scale-rounding error stacks on the quantization error; amax maps to 446.17 / 448 → no clipping
    • round-trip 2.655 % RMS relative, 0.002 % underflow, 0 saturation
    • weight_scale written BF16 [1], matching gorbatjovy's published format (read from its actual safetensors header, not guessed)
    • MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched
  2. Declared it: text_config.ple_embedding_dtype = "float8_e4m3fn".
  3. Renamed layer_types: orca labels its 12 QSA layers qwen_sparse_attention; vLLM accepts only linear_attention / full_attention and picks QSA via indexer_n_heads. ⚠ Verified indexer_n_heads == 4 in BOTH orca and dealignai before renaming — without it the rename silently selects PLAIN attention and serves a subtly wrong model that still looks healthy. Every QSA/indexer key matches dealignai exactly.
  4. .env: FN_MODEL → the converted dir, FN_QUANT → compressed-tensors.

Measured on the live seat

warm decode, conc=1, greedy 300 tok, n=5 median 167.5 tok/s (min 150.0, max 170.4, spread 12.2 %)
MTP k=3 acceptance 60.4 %, mean acceptance length 2.81 (per-pos 80.6/60.8/40.8 %)
KV 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB
on card ~75 GiB; PLE 47.7 GiB pinned host RAM

⚠ Do NOT read 167.5 as a win over dealignai. dealignai's ~121 tok/s in this file came from a different harness/prompt; cross-harness comparison is invalid. ⚠ The 167.5 may itself have been contended — the operator was using the seat around that window — so treat it as a LOWER BOUND, not a clean solo figure. Operator's own session reported 140 tok/s average in real use while the 258K depth probe was running against the same card: an independent, contended floor that agrees with the picture. What IS established is that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did not cost decode speed, which was the standing risk of giving up FP4 tensor-core compute.

Still open on this seat

  • Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context degradation axis). Needs a controlled harness + noise floor.
  • ✅ Deep-prefill probe DONE 2026-09-14: clean to 258,517 tokens. Non-repeating prompts, 6 depths 32K->258K, all 200 OK, and zero allocator OOM/CUBLAS/illegal-memory in the engine log — the detector that caught dealignai's 155K near-miss. Run under real operator load, so a stricter test than solo. Positive control passed (a ~265K prompt got a clean 400 naming the limit). #54919 did not reproduce: 258K prefilled in 28.9 s, ~8,900 tok/s, near-linear. ⚠ The probe's MEMORY column was blind and must not be reused: --kv-cache-memory pins the pool ("skipped memory profiling"), so GPU use is flat vs depth, and before/after nvidia-smi bracketing cannot see a transient mid-prefill spike. Identical readings across an 8x depth range were the tell. Peak-activation headroom remains UNMEASURED. ⭐ Calibration for re-runs: random hex words tokenize at 7.9 tok/word.
  • ⚠⚠ NO LOCAL ROLLBACK. dealignai weights DELETED 2026-09-14 on operator instruction (125 GiB reclaimed). .env.bak-preorca-20260914-023408 still names the old paths but they no longer exist — it is a record, not a revert. Reverting = 126 GiB re-download. The pristine 170 GiB qwen38-flash-next-orcarouter-nvfp4 IS retained (redo the conversion from it; do not delete it without a reason).
  • ✅ Disk reclaimed 2026-09-14: 76 GiB. The 28 non-PLE shards duplicated between the pristine and converted orca dirs are now hardlinked (294G apparent → 218G actual). All 28 verified byte-identical by SHA-256 first, then ln to a temp name + atomic rename over the target — never rm-then-ln, which leaves a window with no file. Done live with the seat serving; it never blinked. ⚠ The two dirs now SHARE INODES. Editing a shared file in place in either dir changes both. Shards are never edited in place, and config.json / model.safetensors.index.json are deliberately NOT shared (the conversion changed them) — but a future session must copy-then-edit, not edit in place.

Other open items

  • cyberprev quant (3rd sec candidate hotdogs/Qwen3.8-27B-abliterated-cyber-preview, bf16 at /tank/aimodels/cyberprev-bf16): restart crashed on a transformers head-count config error in quant-work/.venv — the SAME venv that reached 49/65 on Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path (per services/gen-seat-mixed-quant/README.md), not the venv. Retry that way.
  • sec vs sentinel-r3 quality A/B — both live and gateway-callable; operator to judge.
  • ⚠ Suspect vault entry: fv-gateway/infra-ops-password is 16 chars matching a boot-UUID prefix exactly — possibly malformed. Eyeball.
  • os-nut not installed on the FV OPNsense — with the firewall now on the 5P1000 UPS, a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today.
  • Branch breaker rating + 4-card ammeter reading still open — every power table is arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.