The sentinel-r3 header and two memory notes described mog-sec (Blackfrost
M.O.G.-SEC / Qwentium) as 'a persona system prompt on stock weights'. Its card
is explicit that it is NOT: base_model_relation: finetune on Qwen/Qwen3.8-27B,
a refusal-free offense+defense cybersecurity SFT with YaRN 1M context ('not a
system-prompt sticker on a stock Qwen'). So all three sec-seat candidates are
Qwen3.8-27B SFT finetunes and differ in training focus, not in kind:
mog-sec = broad offense+defense SFT; sentinel-r3 = pentest agent-trajectory SFT;
cyberprev = cyber tool-calling LoRA SFT on an abliterated base.
148 lines
9.8 KiB
Markdown
148 lines
9.8 KiB
Markdown
# 2026-09-14 — fv-ml1 seat reorganization + the gen-large "orca" blocker
|
||
|
||
An all-night GPU-seat overhaul on fv-ml1 after the FV colo recovered. Everything
|
||
here is on the running box; regenerate the authoritative view with
|
||
`scripts/seat-inventory.py` (reads the live containers). Cross-refs:
|
||
[[2026-09-13-flash-next-seat-and-fv-outage]].
|
||
|
||
## Seat topology now (2026-09-14 ~01:40 PT)
|
||
|
||
| GPU | seats |
|
||
|---|---|
|
||
| 0 | `mog-sec` (`sec` :8019, dflash k=7) · **`sentinel-r3`** (`sentinel-r3` :8025, dflash k=7 — NEW) |
|
||
| 1 | **`meromero-charrp`** (`char-rp` :8016, MeroMero-v2-31B dense — restored) · `erp-seat` (`char-rp-fast` :8021) · reward · coder · embed · rerank |
|
||
| 2 | `flash-next` (`gen-large` :8022) — **UP on orcarouter** (PLE converted bf16→FP8; MTP k=3, 60.4% accept) |
|
||
| 3 | **RESERVED scratch** — empty, operator directive; benches/quants/probes only |
|
||
|
||
## What changed tonight
|
||
|
||
1. **flash-next gained MTP k=3.** Campaign in `services/flash-next-mtp-bench/`
|
||
measured MTP a WIN on this hardware (+29/41/27% at k1, +42/52/38% k2, +52/51/34%
|
||
k3 across conc 1/4/8), inverting vLLM's 4×H100 recipe. k=3 deployed (single-user
|
||
→ conc=1 dominates). KV pinned 14→10 GiB (MTP adds 5.08 GiB draft-head weights;
|
||
14 GiB OOMs). Warm decode ~121 tok/s. `stacks/flash-next-seat/compose.yaml`.
|
||
2. **gen consolidation.** All 8 `gen`/`gen-reasoning`/`summarizer`/`summarizer-large`/
|
||
`classifier`/`chat-judge`/`image-judge`/`qwen-image-bench` LiteLLM entries repointed
|
||
to flash-next (:8022); the **27B dense gen seat RETIRED**, freeing 38.4 GB on GPU0.
|
||
⚠ judge aliases now score against different weights — prior scores incomparable.
|
||
3. **char-rp restored to MeroMero-v2-31B.** Was serving a leftover-test RedHatAI 26B
|
||
MoE W4A4; operator wanted the in-house dense-31B heretic W4A16 back. `char-rp-fast`
|
||
(:8021, the 26B MoE) is the deliberate speed tier — the throughput answer, so
|
||
char-rp stays the quality seat. `stacks/meromero-charrp/`.
|
||
4. **Sentinel-R3 A/B + dflash.** `glyphsoftware/sentinel-r3` (proprietary license —
|
||
operator's call) is a REAL SFT pentest finetune vs mog-sec (M.O.G.-SEC), which is
|
||
ALSO a finetune — a refusal-free offense+defense cyber SFT, NOT a persona-on-stock as
|
||
earlier notes claimed. All three sec seats are Qwen3.8-27B finetunes; they differ in focus.
|
||
Served alongside mog-sec for A/B. Then measured dflash vs MTP on Sentinel's
|
||
finetuned body: **dflash 2.40 vs MTP 2.18 mean acceptance length (+11%)**, warm
|
||
decode ~121 tok/s (faster than sec ~102). dflash k=7 cut over. `stacks/sentinel-r3/`.
|
||
⚠ MEASUREMENT LESSON (again): first decode bench read 39 tok/s — a COLD-boot +
|
||
concurrent-contention artifact; warm+isolated it was 121. Operator caught it by
|
||
testing the running `sec` (102 tok/s) as reference.
|
||
|
||
## ✅ THE BLOCKER — RESOLVED 2026-09-14 (no source build needed)
|
||
|
||
gen-large now serves **`orcarouter/Qwen3.8-Flash-Next-Uncensored-NVFP4`** from
|
||
`/tank/aimodels/qwen38-flash-next-orcarouter-nvfp4-plefp8`. Healthy, coherent, MTP k=3.
|
||
|
||
**The earlier diagnosis was right about the symptom and wrong about the cost.** It said the
|
||
only auditable path was cherry-picking a PLE-loader branch onto a from-source Blackwell
|
||
build. Reading the loader in the running nightly showed otherwise:
|
||
|
||
```
|
||
from_quant_config (vllm/models/qwen4_exp/nvidia/ngram_embedding.py:168)
|
||
1. ple_embedding_dtype == "float8_e4m3fn" -> FP8 method <-- BEFORE any type check
|
||
2. quant_config is None -> unquantized
|
||
3. ModelOptMixedPrecisionConfig -> FP8 / unquantized
|
||
4. ModelOptQuantConfigBase + excluded -> unquantized
|
||
5. not isinstance(quant_config, Fp8Config)-> NotImplementedError <-- the blocker
|
||
```
|
||
|
||
Branch 1 is unconditional, and the `NotImplementedError` is **scoped to the PLE embedding
|
||
path only** — experts and dense layers of a compressed-tensors qwen4_exp build load through
|
||
vLLM's ordinary compressed-tensors paths. Falsified directly before doing any work: built
|
||
the real `CompressedTensorsConfig` from orca's own config and called `from_quant_config`
|
||
both ways — as-shipped raises, with the declaration returns `Qwen4ExpPLEFp8EmbeddingMethod`.
|
||
|
||
### What was actually done
|
||
|
||
1. **Converted the PLE table bf16 → FP8.** orca's 128 PLE tensors live in exactly ONE shard
|
||
(`model-00002-of-00017.safetensors`, 95.4 GiB) with **no other tensors in it** — a clean
|
||
split. Converted on GPU3 (reserved scratch) in 139 s into 8 `model-plefp8-*` files.
|
||
- global amax **0.0894**; per-shard outlier ratio only **1.66x**, so one global scale fits
|
||
- scale chosen **exactly representable in bf16** (2.002716e-04) so no scale-rounding error
|
||
stacks on the quantization error; amax maps to **446.17 / 448** → no clipping
|
||
- round-trip **2.655 % RMS relative**, 0.002 % underflow, **0 saturation**
|
||
- `weight_scale` written **BF16 [1]**, matching gorbatjovy's published format (read from
|
||
its actual safetensors header, not guessed)
|
||
- MTP head (31 tensors, BF16) and the 333 vision tensors carried through untouched
|
||
2. **Declared it**: `text_config.ple_embedding_dtype = "float8_e4m3fn"`.
|
||
3. **Renamed `layer_types`**: orca labels its 12 QSA layers `qwen_sparse_attention`; vLLM
|
||
accepts only `linear_attention` / `full_attention` and picks QSA via `indexer_n_heads`.
|
||
⚠ **Verified `indexer_n_heads == 4` in BOTH orca and dealignai before renaming** — without
|
||
it the rename silently selects PLAIN attention and serves a subtly wrong model that still
|
||
looks healthy. Every QSA/indexer key matches dealignai exactly.
|
||
4. `.env`: `FN_MODEL` → the converted dir, `FN_QUANT` → `compressed-tensors`.
|
||
|
||
### Measured on the live seat
|
||
|
||
| | |
|
||
|---|---|
|
||
| warm decode, conc=1, greedy 300 tok, **n=5** | median **167.5 tok/s** (min 150.0, max 170.4, spread 12.2 %) |
|
||
| MTP k=3 | acceptance **60.4 %**, mean acceptance length **2.81** (per-pos 80.6/60.8/40.8 %) |
|
||
| KV | 344,155 tokens @ 262,144 ctx, 1.31x concurrency, pinned 10 GiB |
|
||
| on card | ~75 GiB; PLE 47.7 GiB pinned host RAM |
|
||
|
||
⚠ **Do NOT read 167.5 as a win over dealignai.** dealignai's ~121 tok/s in this file came
|
||
from a different harness/prompt; cross-harness comparison is invalid. ⚠ The 167.5 may itself
|
||
have been contended — the operator was using the seat around that window — so treat it as a
|
||
LOWER BOUND, not a clean solo figure. Operator's own session reported **140 tok/s average**
|
||
in real use while the 258K depth probe was running against the same card: an independent,
|
||
contended floor that agrees with the picture. What IS established is
|
||
that weight-only experts (orca is W8 weight-only attn + W4 weight-only experts) did **not**
|
||
cost decode speed, which was the standing risk of giving up FP4 tensor-core compute.
|
||
|
||
### Still open on this seat
|
||
|
||
- Quality A/B orca vs dealignai (the actual reason for the swap — the W4A4 long-context
|
||
degradation axis). Needs a controlled harness + noise floor.
|
||
- ✅ **Deep-prefill probe DONE 2026-09-14: clean to 258,517 tokens.** Non-repeating prompts,
|
||
6 depths 32K->258K, all 200 OK, and **zero allocator OOM/CUBLAS/illegal-memory in the engine
|
||
log** — the detector that caught dealignai's 155K near-miss. Run under real operator load, so
|
||
a stricter test than solo. Positive control passed (a ~265K prompt got a clean 400 naming the
|
||
limit). #54919 did not reproduce: 258K prefilled in 28.9 s, ~8,900 tok/s, near-linear.
|
||
⚠ The probe's MEMORY column was blind and must not be reused: `--kv-cache-memory` pins the
|
||
pool ("skipped memory profiling"), so GPU use is flat vs depth, and before/after `nvidia-smi`
|
||
bracketing cannot see a transient mid-prefill spike. Identical readings across an 8x depth
|
||
range were the tell. Peak-activation headroom remains UNMEASURED.
|
||
⭐ Calibration for re-runs: random hex words tokenize at **7.9 tok/word**.
|
||
- ⚠⚠ **NO LOCAL ROLLBACK.** dealignai weights DELETED 2026-09-14 on operator instruction
|
||
(125 GiB reclaimed). `.env.bak-preorca-20260914-023408` still names the old paths but they
|
||
no longer exist — it is a record, not a revert. Reverting = 126 GiB re-download.
|
||
The pristine 170 GiB `qwen38-flash-next-orcarouter-nvfp4` IS retained (redo the conversion
|
||
from it; do not delete it without a reason).
|
||
- ✅ **Disk reclaimed 2026-09-14: 76 GiB.** The 28 non-PLE shards duplicated between the
|
||
pristine and converted orca dirs are now **hardlinked** (294G apparent → 218G actual).
|
||
All 28 verified **byte-identical by SHA-256** first, then `ln` to a temp name + atomic
|
||
`rename` over the target — never rm-then-ln, which leaves a window with no file. Done
|
||
live with the seat serving; it never blinked.
|
||
⚠ **The two dirs now SHARE INODES.** Editing a shared file *in place* in either dir
|
||
changes both. Shards are never edited in place, and `config.json` /
|
||
`model.safetensors.index.json` are deliberately NOT shared (the conversion changed
|
||
them) — but a future session must copy-then-edit, not edit in place.
|
||
|
||
## Other open items
|
||
|
||
- **cyberprev quant** (3rd sec candidate `hotdogs/Qwen3.8-27B-abliterated-cyber-preview`,
|
||
bf16 at `/tank/aimodels/cyberprev-bf16`): restart crashed on a transformers
|
||
head-count config error in `quant-work/.venv` — the SAME venv that reached 49/65 on
|
||
Sept 11, so the original run likely used the canonical vllm-image+llmcompressor path
|
||
(per `services/gen-seat-mixed-quant/README.md`), not the venv. Retry that way.
|
||
- **sec vs sentinel-r3 quality A/B** — both live and gateway-callable; operator to judge.
|
||
- ⚠ **Suspect vault entry**: `fv-gateway/infra-ops-password` is 16 chars matching a
|
||
boot-UUID prefix exactly — possibly malformed. Eyeball.
|
||
- **os-nut not installed** on the FV OPNsense — with the firewall now on the 5P1000 UPS,
|
||
a NUT transfer-to-battery event is a direct "breaker tripped" alarm nobody gets today.
|
||
- **Branch breaker rating + 4-card ammeter reading** still open — every power table is
|
||
arithmetic on an ESTIMATED ~300 W platform draw. The ammeter converts it to fact.
|