memory: granite spike = mechanical-green ONLY, efficacy not validated by design (mtf-dev confirm)
The granite-8b harness spike proved the TRL SFT->DPO->eval seam (incl. the in-loop HoldoutEvaluator base-vs-adapter leg) runs end-to-end, but used a 12-row/12-pair synthetic writing fixture — NOT the E-RP corpus — so the ~0 anti-slop delta (-0.002) is the expected null, not an efficacy signal. Adapter reaped; nothing to A/B. Real behaviour-shift efficacy is a T1-run question. Sharpen both the T1 in-flight bullet and the Recent-decisions entry so 'green' no longer reads as efficacy-validated.
This commit is contained in:
+15
-4
@@ -121,9 +121,11 @@ _As of 2026-07-02:_
|
||||
|
||||
- **T1 (mtf-dev qwopus writing-LoRA) IN PROGRESS.** Attention-only LoRA (full-attn q/k/v/o
|
||||
+ the bf16 GDN `in_proj_qkv`/`out_proj`); serve via swappable-LoRA-on-NVFP4 (test QUEUED,
|
||||
gated on the first T1 adapter) or merge+requant. Trainer harness seam PROVEN via a
|
||||
granite-8b spike (green 2026-07-02). NVFP4 quant-structure + serve-path facts in
|
||||
`reference_gen_qwopus_122b`.
|
||||
gated on the first T1 adapter) or merge+requant. Trainer harness seam MECHANICALLY
|
||||
proven via a granite-8b spike (2026-07-02): the pipeline runs end-to-end, but efficacy
|
||||
is NOT validated (mechanics-only synthetic fixture, by design; a granite adapter
|
||||
wouldn't transfer to qwopus regardless). Real behaviour-shift efficacy is a T1-run
|
||||
question. NVFP4 quant-structure + serve-path facts in `reference_gen_qwopus_122b`.
|
||||
|
||||
- **Worldtree #332 scoped-log view + tunnel = STANDING ASSET.** `wt_gateway_logs` view +
|
||||
`wt_readonly` role on the litellm DB + `wt-db-tunnel` systemd on corviduo-dev, for
|
||||
@@ -176,7 +178,16 @@ _As of 2026-07-02:_
|
||||
with comfy-dev), staged granite-4.1-8b bf16 to `irv-ml1:/home/lkraven`, mtf-dev's TRL
|
||||
SFT→DPO→eval seam proved end-to-end (DPO genuinely learned, 0.833 acc; anti-slop ~0 = the
|
||||
expected null on clean-writing granite-instruct); comfyui restored healthy. De-risks T1's
|
||||
trainer harness ahead of the real qwopus train.
|
||||
trainer harness ahead of the real qwopus train. **MECHANICAL green ONLY — efficacy NOT
|
||||
validated, by design** (mtf-dev confirm 2026-07-02): the run used a 12-row/12-pair
|
||||
SYNTHETIC writing fixture (clean-vs-sloppy generic prompts), NOT the E-RP corpus; rank 8,
|
||||
1 epoch, ~2 SFT + ~6 DPO steps. The in-loop base-vs-adapter check (HoldoutEvaluator)
|
||||
already ran and returned the expected null (anti_slop_improvement −0.002); the adapter was
|
||||
then reaped (gone), so there is nothing to A/B — and on a mechanics-only synthetic adapter
|
||||
an A/B would only reconfirm the null. Real efficacy = the T1 run (real recipe + E-RP data
|
||||
on qwopus). Open (mtf-dev routing to operator): whether to insert an intermediate
|
||||
*real-efficacy* granite spike before T1 — weak proxy (granite arch ≠ qwopus, + granite is
|
||||
censored vs qwopus abliterated), leaning defer-to-T1.
|
||||
|
||||
- `[2026-07-01]` **Worldtree #332 embed-recall diagnosed + scoped-log view/tunnel
|
||||
provisioned + fix verified.** The persona-recitation gate re-embedded persona segments
|
||||
|
||||
Reference in New Issue
Block a user