Files
esh-pfi-infrastructure/persistent-memory.d/2026-09-14-fv-seat-reorg-and-orca-blocker.md
T
vh 3906c6842c docs: correct sec-seat lineage — M.O.G.-SEC/mog-sec is an offense+defense SFT finetune, not a persona-on-stock
The sentinel-r3 header and two memory notes described mog-sec (Blackfrost
M.O.G.-SEC / Qwentium) as 'a persona system prompt on stock weights'. Its card
is explicit that it is NOT: base_model_relation: finetune on Qwen/Qwen3.8-27B,
a refusal-free offense+defense cybersecurity SFT with YaRN 1M context ('not a
system-prompt sticker on a stock Qwen'). So all three sec-seat candidates are
Qwen3.8-27B SFT finetunes and differ in training focus, not in kind:
mog-sec = broad offense+defense SFT; sentinel-r3 = pentest agent-trajectory SFT;
cyberprev = cyber tool-calling LoRA SFT on an abliterated base.
2026-09-14 07:31:23 -07:00

148 lines
9.8 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# 2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker
An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything
here is on the running box; regenerate the authoritative view with
`scripts/seat-inventory.py` (reads the live containers). Cross-refs:
[[2026-09-13-flash-next-seat-and-fv-outage]].
## Seat topology now (2026-09-14 ~01:40 PT)
| GPU | seats |
|---|---|
| 0 | `mog-sec` (`sec` :8019, dflash k=7) · **`sentinel-r3`** (`sentinel-r3` :8025, dflash k=7 — NEW) |
| 1 | **`meromero-charrp`** (`char-rp` :8016, MeroMero-v2-31B dense — restored) · `erp-seat` (`char-rp-fast` :8021) · reward · coder · embed · rerank |
| 2 | `flash-next` (`gen-large` :8022) — **UP on orcarouter** (PLE converted bf16→FP8; MTP k=3, 60.4% accept) |
| 3 | **RESERVED scratch** — empty, operator directive; benches/quants/probes only |
## What changed tonight
1. **flash-next gained MTP k=3.** Campaign in `services/flash-next-mtp-bench/`
measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34%
k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user
→ conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights;
14 GiB OOMs). Warm decode ~121 tok/s. `stacks/flash-next-seat/compose.yaml`.
2. **gen consolidation.** All 8 `gen`/`gen-reasoning`/`summarizer`/`summarizer-large`/
`classifier`/`chat-judge`/`image-judge`/`qwen-image-bench` LiteLLM entries repointed
to flash-next (:8022); the **27B dense gen seat RETIRED**, freeing 38.4 GB on GPU0.
⚠ judge aliases now score against different weights — prior scores incomparable.
3. **char-rp restored to MeroMero-v2-31B.** Was serving a leftover-test RedHatAI 26B
MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back. `char-rp-fast`
(:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so
char-rp stays the quality seat. `stacks/meromero-charrp/`.
4. **Sentinel-R3 A/B + dflash.** `glyphsoftware/sentinel-r3` (proprietary license —
operator's call) is a REAL SFT pentest finetune vs mog-sec (M.O.G.-SEC), which is
ALSO a finetune — a refusal-free offense+defense cyber SFT, NOT a persona-on-stock as
earlier notes claimed. All three sec seats are Qwen3.8-27B finetunes; they differ in focus.
Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's
finetuned body: **dflash 2.40 vs MTP 2.18 mean acceptance length (+11%)**, warm
decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over. `stacks/sentinel-r3/`.
⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot +
concurrent-contention artifact; warm+isolated it was 121. Operator caught it by
testing the running `sec` (102 tok/s) as reference.
## ✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed)
gen-large now serves **`orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4`** from
`/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8`. Healthy, coherent, MTP k=3.
**The earlier diagnosis was right about the symptom and wrong about the cost.** It said the
only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell
build. Reading the loader in the running nightly showed otherwise:
```
from_quant_config (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168)
1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method <-- BEFORE any type check
2. quant_config is None -> unquantized
3. ModelOptMixedPrecisionConfig -> FP8 / unquantized
4. ModelOptQuantConfigBase + excluded -> unquantized
5. not isinstance(quant_config, Fp8Config)-> NotImplementedError <-- the blocker
```
Branch 1 is unconditional, and the `NotImplementedError` is **scoped to the PLE embedding
path only** — experts and dense layers of a compressed-tensors qwen4_exp build load through
vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built
the real `CompressedTensorsConfig` from orca's own config and called `from_quant_config`
both ways — as-shipped raises, with the declaration returns `Qwen4ExpPLEFp8EmbeddingMethod`.
### What was actually done
1. **Converted the PLE table bf16 → FP8.** orca's 128 PLE tensors live in exactly ONE shard
(`model-00002-of-00017.safetensors`, 95.4 GiB) with **no other tensors in it** — a clean
split. Converted on GPU3 (reserved scratch) in 139 s into 8 `model-plefp8-*` files.
- global amax **0.0894**; per-shard outlier ratio only **1.66x**, so one global scale fits
- scale chosen **exactly representable in bf16** (2.002716e-04) so no scale-rounding error
stacks on the quantization error; amax maps to **446.17 / 448** → no clipping
- round-trip **2.655 % RMS relative**, 0.002 % underflow, **0 saturation**
- `weight_scale` written **BF16 [1]**, matching gorbatjovy's published format (read from
its actual safetensors header, not guessed)
- MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched
2. **Declared it**: `text_config.ple_embedding_dtype = "float8_e4m3fn"`.
3. **Renamed `layer_types`**: orca labels its 12 QSA layers `qwen_sparse_attention`; vLLM
accepts only `linear_attention` / `full_attention` and picks QSA via `indexer_n_heads`.
⚠ **Verified `indexer_n_heads == 4` in BOTH orca and dealignai before renaming** — without
it the rename silently selects PLAIN attention and serves a subtly wrong model that still
looks healthy. Every QSA/indexer key matches dealignai exactly.
4. `.env`: `FN_MODEL` → the converted dir, `FN_QUANT` → `compressed-tensors`.
### Measured on the live seat
| | |
|---|---|
| warm decode, conc=1, greedy 300 tok, **n=5** | median **167.5 tok/s** (min 150.0, max 170.4, spread 12.2 %) |
| MTP k=3 | acceptance **60.4 %**, mean acceptance length **2.81** (per-pos 80.6/60.8/40.8 %) |
| KV | 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB |
| on card | ~75 GiB; PLE 47.7 GiB pinned host RAM |
⚠ **Do NOT read 167.5 as a win over dealignai.** dealignai's ~121 tok/s in this file came
from a different harness/prompt; cross-harness comparison is invalid. ⚠ The 167.5 may itself
have been contended — the operator was using the seat around that window — so treat it as a
LOWER BOUND, not a clean solo figure. Operator's own session reported **140 tok/s average**
in real use while the 258K depth probe was running against the same card: an independent,
contended floor that agrees with the picture. What IS established is
that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did **not**
cost decode speed, which was the standing risk of giving up FP4 tensor-core compute.
### Still open on this seat
- Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context
degradation axis). Needs a controlled harness + noise floor.
- ✅ **Deep-prefill probe DONE 2026-09-14: clean to 258,517 tokens.** Non-repeating prompts,
6 depths 32K->258K, all 200 OK, and **zero allocator OOM/CUBLAS/illegal-memory in the engine
log** — the detector that caught dealignai's 155K near-miss. Run under real operator load, so
a stricter test than solo. Positive control passed (a ~265K prompt got a clean 400 naming the
limit). #54919 did not reproduce: 258K prefilled in 28.9 s, ~8,900 tok/s, near-linear.
⚠ The probe's MEMORY column was blind and must not be reused: `--kv-cache-memory` pins the
pool ("skipped memory profiling"), so GPU use is flat vs depth, and before/after `nvidia-smi`
bracketing cannot see a transient mid-prefill spike. Identical readings across an 8x depth
range were the tell. Peak-activation headroom remains UNMEASURED.
⭐ Calibration for re-runs: random hex words tokenize at **7.9 tok/word**.
- ⚠⚠ **NO LOCAL ROLLBACK.** dealignai weights DELETED 2026-09-14 on operator instruction
(125 GiB reclaimed). `.env.bak-preorca-20260914-023408` still names the old paths but they
no longer exist — it is a record, not a revert. Reverting = 126 GiB re-download.
The pristine 170 GiB `qwen38-flash-next-orcarouter-nvfp4` IS retained (redo the conversion
from it; do not delete it without a reason).
- ✅ **Disk reclaimed 2026-09-14: 76 GiB.** The 28 non-PLE shards duplicated between the
pristine and converted orca dirs are now **hardlinked** (294G apparent → 218G actual).
All 28 verified **byte-identical by SHA-256** first, then `ln` to a temp name + atomic
`rename` over the target — never rm-then-ln, which leaves a window with no file. Done
live with the seat serving; it never blinked.
⚠ **The two dirs now SHARE INODES.** Editing a shared file *in place* in either dir
changes both. Shards are never edited in place, and `config.json` /
`model.safetensors.index.json` are deliberately NOT shared (the conversion changed
them) — but a future session must copy-then-edit, not edit in place.
## Other open items
- **cyberprev quant** (3rd sec candidate `hotdogs/Qwen3.8-27B-abliterated-cyber-preview`,
bf16 at `/tank/aimodels/cyberprev-bf16`): restart crashed on a transformers
head-count config error in `quant-work/.venv` — the SAME venv that reached 49/65 on
Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path
(per `services/gen-seat-mixed-quant/README.md`), not the venv. Retry that way.
- **sec vs sentinel-r3 quality A/B** — both live and gateway-callable; operator to judge.
- ⚠ **Suspect vault entry**: `fv-gateway/infra-ops-password` is 16 chars matching a
boot-UUID prefix exactly — possibly malformed. Eyeball.
- **os-nut not installed** on the FV OPNsense — with the firewall now on the 5P1000 UPS,
a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today.
- **Branch breaker rating + 4-card ammeter reading** still open — every power table is
arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.