memory: snapshot — run 3 gated DO-NOT-SERVE, run 3c held on a tripped breaker
Run 3 trained, gated and dispositioned do-not-serve on a measured 44pp self-harm guardrail regression that its own preregistered rule passed -- a pooled preserve-list test cannot see a single-axis collapse. Run 3c (lr 20x cut, single variable) launched, killed by an Anaheim power-breaker trip at step 80, relaunched, then stopped by the operator at step 22 pending a weekend power triage. Also captured: the corpus mix was specified in a unit the optimiser never sees (45.8% dialogue by context, 24.2% by loss); the dose-response says benefit and damage are one direction in weight space, so the merge-back measures the problem rather than fixing it; four guests including the storage SPOF had onboot unset and never came back from the outage, now fixed with dependency ordering; and a transport failure that enters a measurement as a value looks like whatever you hoped to find -- which found a live defect in another agent's instrument an hour after it was reported. Auto-archived 8 entries to archival-memory.md (Recent decisions: 8, Tried and abandoned: 0); 4 held back on open deferred-work pointers.
This commit is contained in:
@@ -0,0 +1,62 @@
|
||||
# `[2026-08-27]` The dose-response says benefit and damage are ONE direction in weight space
|
||||
|
||||
vLLM **cannot LoRA-serve Gemma-4-26B-A4B at all** — it is an MoE and the LoRA manager has no
|
||||
expert mapping (`AttributeError: To support LoRA for MoE model, 'get_expert_mapping' must be
|
||||
implemented`, `vllm/lora/utils.py:398`, v0.24.0). Found by trying it: one container start,
|
||||
which is exactly what playbook §3.10 exists to buy.
|
||||
|
||||
So each scale point was **pre-merged** instead — `merge_and_unload` with the adapter's alpha
|
||||
pre-scaled, which reproduces exactly what a LoRA at scale s would serve, on the same serving
|
||||
path as the gate. Artifacts held at `/tank/erp-tune/serve/merged-run03{,-s075,-s050,-s025}`.
|
||||
|
||||
scale attractor gain kept noise@31 damage prose@40 cost
|
||||
s=0.50 0.9684 0.0% 1.0000 0.00pp 0.00pp
|
||||
s=0.75 0.8966 42.0% 0.9688 3.12pp 6.25pp
|
||||
s=1.00 0.8049 100.0% 0.5938 40.62pp 15.62pp
|
||||
|
||||
first 42% of the gain costs 3.12pp of noise coherence
|
||||
last 58% of the gain costs 37.50pp more -> 12x worse per point
|
||||
|
||||
**Every axis is monotone in scale — attractor, distance, marker density, RP length, both
|
||||
coherence measures, even truncation count. No knee anywhere.**
|
||||
|
||||
## The conclusion, which is stronger than an exchange rate
|
||||
|
||||
An exchange rate says the trade is expensive. This says **there is no trade to make**: the
|
||||
adapter learned ONE direction and everything rides it. brokkr's framing, worth keeping:
|
||||
**scaling moves you along the direction the adapter already learned; it cannot give you a
|
||||
different one.**
|
||||
|
||||
That converts the merge-back from a candidate fix into a **measurement of the problem**.
|
||||
zerofata's MeroMero v1 ships the merge-back as its answer, so it was worth testing — and
|
||||
testing it is what proved it cannot be the answer here.
|
||||
|
||||
## The T4 mechanism is NOT register capture
|
||||
|
||||
pad type base tuned
|
||||
prose 1.0000 flat mild knee, onset ~1,600 tok
|
||||
noise 1.0000 flat 1.0000 -> 0.5938 <- SEVERE
|
||||
|
||||
The prose-gradient "wedge" predicted **prose** should be the worst case. Instead prose is the
|
||||
mild axis and **shuffled tokens are where it collapses** — a continuation prior has nothing to
|
||||
continue in shuffled tokens. Every miss is `wrong-name`: correct format, on task, wrong answer.
|
||||
**State mis-tracking, not register capture.**
|
||||
|
||||
brokkr: *"The tune did not teach it to continue prose — it destroyed its ability to skip text
|
||||
that is not worth reading."*
|
||||
|
||||
## ⚠ INSTRUMENT FLOOR — temperature 0 is NOT deterministic under concurrency
|
||||
|
||||
same seat, same items, temperature 0, --jobs 8
|
||||
prose reps 31 0.9375 0.9688 0.9688 spread 3.12pp
|
||||
noise reps 24 0.7500 0.7188 0.7188 spread 3.12pp
|
||||
|
||||
vLLM's continuous batching is not numerically invariant to batch composition. **Any eval delta
|
||||
under ~3.1pp at n=64/jobs=8 is inside the floor.** Only `--jobs 1` shrinks it; more n averages
|
||||
over it. Also in auto-memory as
|
||||
`reference_vllm_temp0_not_deterministic_under_jobs`.
|
||||
|
||||
Corollary: it retro-closes the window-2 concurrency worry — measured dirty-vs-clean divergence
|
||||
was 1.56pp, comfortably inside this floor.
|
||||
|
||||
See [[2026-08-27-mix-specified-in-the-wrong-unit]] for the wedge this displaced.
|
||||
Reference in New Issue
Block a user