memory: retract the MTP-head degeneration hypothesis; n=1 was never evidence
Operator ruling: the multi-turn degeneration lives in the un-fixed vLLM, not in the weights. The hypothesis that sec's stock-graft MTP head causes it is withdrawn. Two failures produced it. First, a false dichotomy treated as a deduction: having verified gen and sec run an identical engine, I concluded config was eliminated and therefore the weights were responsible. That does not follow. An engine bug present in both seats is not exonerated by the seats being identical; it only means the engine cannot explain a difference between them. It can still explain the failure. Second, and more instructive, the difference being explained may not exist. The premise was a single operator observation made during a session with many concurrent changes. That cannot carry a causal claim, and it became the load-bearing support for a root-cause narrative it could not hold. The same caveat now attaches to the coherent-to-10k observation on the new build: same n, same uncontrolled conditions, opposite direction. The comparison is weak at both ends, so the file no longer presents either sighting as a result. What survives as measured fact is unchanged and still recorded: sec's MTP head is byte-identical to the uncensored base across all 15 tensors, gen's was abliterated in-band, and acceptance differs slightly. None of that is shown to cause degeneration. Adds the generalisable lesson: an observation made while many things are changing cannot support a causal conclusion. It is the inverse of the warning already in the gen-seat compose file, which guards against trusting a negative result from a synthetic probe; this guards against trusting a positive sighting from an uncontrolled session.
This commit is contained in:
@@ -70,31 +70,60 @@ plausible (the drafter reads quantized hidden states at its five taps). Second c
|
||||
prompt distribution (ours general-purpose, theirs GSM8K/MATH/HumanEval/MBPP/MT-Bench).
|
||||
**Neither is confirmed.** Settling it needs a BF16 target seat (~56 GB) — a real GPU window.
|
||||
|
||||
## 🔶 HYPOTHESIS — why sec degenerated and gen did not
|
||||
## ❌ RETRACTED — the "MTP head mismatch causes the degeneration" hypothesis
|
||||
|
||||
Operator confirmed **sec degenerates at ~2k tokens as served**, gen does not, on identical
|
||||
engines. Config was eliminated as a variable:
|
||||
**Operator ruling, 2026-08-22: this hypothesis is WRONG. The degeneration lives in the un-fixed
|
||||
vLLM, not in the weights.** Recorded here rather than deleted, because it was reasoned to
|
||||
confidently enough that a future session could re-derive it.
|
||||
|
||||
- **Same vLLM image ID** `sha256:bd3236cff208…`, same live version `0.27.2rc1.dev150+g311b3513a`
|
||||
read from inside both running processes (tag equality alone would not prove this).
|
||||
- Same `qwen3_5_mtp` k=3, same `--enable-prefix-caching`, same `fp8` KV, same `float32` mamba
|
||||
cache. Only deltas were `gpu-memory-utilization` 0.43 vs 0.44 and the served name.
|
||||
**Two independent failures produced it, and the second is the instructive one:**
|
||||
|
||||
**The standing hypothesis is the MTP head.** Per our own provenance (verified with
|
||||
`compare_mtp_head.py`): sec's head is **byte-identical to `qwen38-27b-uncensored-bf16`, all 15
|
||||
tensors** — a stock head against a security-finetuned body, because the
|
||||
`Qwen3_5ForConditionalGeneration` wrapper never loads the head so the finetuning could not reach
|
||||
it. gen's orcarouter head was **abliterated in-band by the author**, matched to its body.
|
||||
Acceptance corroborates (gen 58.4% vs sec 55.9%).
|
||||
1. **I treated a false dichotomy as a deduction.** Having verified gen and sec run an identical
|
||||
engine (same image ID `sha256:bd3236cff208…`, same live version
|
||||
`0.27.2rc1.dev150+g311b3513a` read from inside both processes, same flags bar
|
||||
`gpu-memory-utilization` 0.43 vs 0.44), I concluded "config is eliminated, therefore it is the
|
||||
weights." That does not follow. **An engine bug present in BOTH seats is not exonerated by the
|
||||
two seats being identical** — it just means the engine cannot explain a *difference*. It can
|
||||
still explain the *failure*.
|
||||
2. **The difference I was explaining may not exist.** The premise was a single operator
|
||||
observation of sec degenerating at ~2k, made during a session with many concurrent changes.
|
||||
**n=1 under heavy concurrent modification is not evidence** — see the meta-lesson below.
|
||||
|
||||
**⚠ This is consistent with everything measured but is NOT proven.** Nobody has shown the head
|
||||
mismatch *causes* the degeneration.
|
||||
**What survives as fact** (measured, still true, just not causal): sec's MTP head *is*
|
||||
byte-identical to `qwen38-27b-uncensored-bf16` across all 15 tensors — a stock head on a
|
||||
security-finetuned body, because the `Qwen3_5ForConditionalGeneration` wrapper never loads the
|
||||
head, so the finetuning could not reach it. gen's orcarouter head *was* abliterated in-band by
|
||||
its author. Acceptance differs slightly (gen 58.4%, sec 55.9%). **All true. None of it shown to
|
||||
cause multi-turn degeneration.**
|
||||
|
||||
## ⚠️ CONFOUNDED — what fixed sec is not yet known
|
||||
**Current standing explanation: the degeneration is an engine bug in the un-fixed vLLM.** Both
|
||||
production seats run `311b3513`, which is **172 commits behind GDN spec-decode fix #53077**
|
||||
(merged 2026-08-20). `#51113` is present in that build and is therefore **necessary but
|
||||
insufficient** on its own.
|
||||
|
||||
## ⭐⭐ META-LESSON — n=1 during a busy session is not evidence
|
||||
|
||||
The operator's own framing, and it generalises past this incident: **an observation made while
|
||||
many things are being changed at once cannot carry a causal claim, no matter how confidently it
|
||||
is reported.** Tonight that single observation became the load-bearing premise for a weights-side
|
||||
hypothesis, a root-cause narrative, and very nearly a recommendation.
|
||||
|
||||
This is the same failure the gen-seat compose file already warns about in different words — *"a
|
||||
passing probe is NOT sufficient evidence"* — inverted. That note guards against trusting a
|
||||
**negative** result from a synthetic test. This one guards against trusting a **positive**
|
||||
sighting from an uncontrolled session. Both reduce to: **hold the system still, or do not draw
|
||||
causal conclusions from it.**
|
||||
|
||||
Applies equally to the "coherent to 10k" observation below — same n, same conditions, opposite
|
||||
direction. Neither observation is worth more than the other.
|
||||
|
||||
## ⚠️ CONFOUNDED — and the "before" state is itself unreliable
|
||||
|
||||
sec now runs DFlash2 on a newer build and the operator reports **coherent to 10k tokens with
|
||||
adversarial nonsense prompts**, where it degenerated at 2k before. **Two variables changed at
|
||||
once:**
|
||||
adversarial nonsense prompts**. ⚠ Treat this the same way as the 2k sighting it is being compared
|
||||
against: **n=1, uncontrolled session, not evidence.** The comparison is weak on *both* ends.
|
||||
|
||||
**Two variables changed at once:**
|
||||
|
||||
1. **Engine**: `311b3513` → `e9d1398d`, **+259 commits, `behind_by=0`** (a strict superset),
|
||||
including GDN spec-decode fix **#53077** (merged 2026-08-20) that production is **172 commits
|
||||
|
||||
Reference in New Issue
Block a user