1.2 KiB
Bonsai ternary spike on fv-ml1 GPU 3 (Prime via brokkr): at the 275 W cap, PQ2_0 is 1.93x Q4_K_XL at N=1 but only 1.06x at N=8 (PTQ1_0 1.24x), recovering to ~1.37x at N=16.
[2026-09-28] Bonsai ternary spike on fv-ml1 GPU 3 (Prime via brokkr): at the 275 W cap, PQ2_0 is 1.93x Q4_K_XL at N=1 but only 1.06x at N=8 (PTQ1_0 1.24x), recovering to ~1.37x at N=16. Positive control passed on tg128 (+1.1%). nvidia-smi's sw_power_cap flag never fires on this card, so "at cap" is judged from board draw. Follow-up (brokkr): moving PQ2_0 onto MMQ from batch 6, like Q4_K (mmvq.cu, ne11 <= 5), lifts N=8 to 1.21x, so the dip was partly a kernel threshold. The residual gap is not power. Runs are in fv-ml1:/tank/spikes/bonsai-2026-09-28/runs{,-mmvq5}/ and on the Booth. The build image was removed; Dockerfile.build recreates it. Acquired for keeps (Prime, 1540): GGUF PQ2_0/PTQ1_0/mmproj-Q8_0 + the fork source pin in /tank/aimodels/llm/prism-ml_Ternary-Bonsai-2-27B-gguf/ (runtime/), the MLX 2-bit pack in /tank/aimodels/mlx/, 34/34 hash-verified. Weights single copy (no snapshots, /tank not in restic); the fork pin also at /mnt/smithy/runtime-pins/prism-llama.cpp-87268f77/ (brokkr).