d710e56aca
Split the 53 over-threshold dated log entries (Recent decisions, Tried and abandoned) into per-entry persistent-memory.d/<slug>.md detail files, leaving one-line pointers in the index; the 10 short entries stay inline. Startup index drops 60,527 -> 25,256 bytes (492 -> 270 lines); entry bodies move verbatim to on-demand detail files, so a fresh session loads ~25 KB instead of ~60 KB and pulls a detail file only when its pointer is relevant. Top matter (Repo purpose, Tools & conventions, Current state) is unchanged; both archival back-references preserved. CLAUDE.md persistent-memory section now documents the index<->detail read discipline (read the index, pull details on demand, never bulk-read the dir, commit both together). Auto-archival still held every dated entry back (all <30 days old); the July burst begins aging past the 30-day guard ~2026-07-31.
1.2 KiB
1.2 KiB
[2026-07-14]NVFP4 quant chase RESOLVED (gibberish) + PIVOTED to modelopt for MTP. One ~40-min GPU0 window. Root-caused the!!!!to the quant NAMESPACE (text-onlyAutoModelForCausalLM→model.layers.*keys; vLLM serves onlyQwen3_5ForConditionalGeneration, which needsmodel.language_model.*) — found from config diffs + vLLM source with ZERO GPU time; fixed by loading asAutoModelForImageTextToText. NVFP4 now serves COHERENT (validated greedy). BUT base NVFP4 ≈53 tok/s ≈ GGUF's 59.5 at batch-1 (no single-stream win) AND MTP = 0% acceptance on compressed-tensors (bf16 mtp head only loads on the modelopt format). Operator chose to pursue a modelopt-format re-quant (the only path to the 2-4× MTP goal; AEON-proven on this exact Qwen3.6-27B arch). Scoped + de-risked: AEON/tank/aimodels/qwen36-27b-aeon-nvfp4= the modelopt reference (quant_method modelopt, 1967 tensors, 15 bf16 mtp keys identical to graft); nvidia-modelopt 0.45.0 installs +mtq.quantize/NVFP4_DEFAULT_CFG/export_hf_checkpointAPI confirmed; pipeline unchanged except swap llm-compressor→modelopt. Seats restored; char-rp-reasoning stays GGUF. Full plan in Current state ★ section.