memory: granite spike = mechanical-green ONLY, efficacy not validated by design (mtf-dev confirm)

The granite-8b harness spike proved the TRL SFT->DPO->eval seam (incl. the
in-loop HoldoutEvaluator base-vs-adapter leg) runs end-to-end, but used a
12-row/12-pair synthetic writing fixture — NOT the E-RP corpus — so the
~0 anti-slop delta (-0.002) is the expected null, not an efficacy signal.
Adapter reaped; nothing to A/B. Real behaviour-shift efficacy is a T1-run
question. Sharpen both the T1 in-flight bullet and the Recent-decisions
entry so 'green' no longer reads as efficacy-validated.
This commit is contained in:
2026-07-02 10:25:46 -07:00
parent 5b52673b75
commit 7fdda2de53
+15 -4
View File
@@ -121,9 +121,11 @@ _As of 2026-07-02:_
- **T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS.** Attention-only LoRA (full-attn q/k/v/o
+ the bf16 GDN `in_proj_qkv`/`out_proj`); serve via swappable-LoRA-on-NVFP4 (test QUEUED,
gated on the first T1 adapter) or merge+requant. Trainer harness seam PROVEN via a
granite-8b spike (green 2026-07-02). NVFP4 quant-structure + serve-path facts in
`reference_gen_qwopus_122b`.
gated on the first T1 adapter) or merge+requant. Trainer harness seam MECHANICALLY
proven via a granite-8b spike (2026-07-02): the pipeline runs end-to-end, but efficacy
is NOT validated (mechanics-only synthetic fixture, by design; a granite adapter
wouldn't transfer to qwopus regardless). Real behaviour-shift efficacy is a T1-run
question. NVFP4 quant-structure + serve-path facts in `reference_gen_qwopus_122b`.
- **Worldtree #332 scoped-log view + tunnel = STANDING ASSET.** `wt_gateway_logs` view +
`wt_readonly` role on the litellm DB + `wt-db-tunnel` systemd on corviduo-dev, for
@@ -176,7 +178,16 @@ _As of 2026-07-02:_
with comfy-dev), staged granite-4.1-8b bf16 to `irv-ml1:/home/lkraven`, mtf-dev's TRL
SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the
expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's
trainer harness ahead of the real qwopus train.
trainer harness ahead of the real qwopus train. **MECHANICAL green ONLY — efficacy NOT
validated, by design** (mtf-dev confirm 2026-07-02): the run used a 12-row/12-pair
SYNTHETIC writing fixture (clean-vs-sloppy generic prompts), NOT the E-RP corpus; rank 8,
1 epoch, ~2 SFT + ~6 DPO steps. The in-loop base-vs-adapter check (HoldoutEvaluator)
already ran and returned the expected null (anti_slop_improvement 0.002); the adapter was
then reaped (gone), so there is nothing to A/B — and on a mechanics-only synthetic adapter
an A/B would only reconfirm the null. Real efficacy = the T1 run (real recipe + E-RP data
on qwopus). Open (mtf-dev routing to operator): whether to insert an intermediate
*real-efficacy* granite spike before T1 — weak proxy (granite arch ≠ qwopus, + granite is
censored vs qwopus abliterated), leaning defer-to-T1.
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel
provisioned + fix verified.** The persona-recitation gate re-embedded persona segments