feat(char-rp-gguf): max context — char-rp 96K, char-rp-reasoning 40K, q8_0 KV

Raise both RP seats to near-max context within the GPU0 budget using q8_0 KV cache
(near-lossless 8-bit, ~2x context/GB, flash-attn-backed). Verified coherent on both
(no Qwen KV-quant gibberish) at 64/50 tok/s.

- char-rp (Magidonia): 16K -> 96K (native 128K; 128K would starve the reasoning seat).
- char-rp-reasoning (QwQ): 16K -> 40960 (QwQ native max; beyond needs YaRN).
- kv_unified=true -> a single conversation gets the full n_ctx (slots share the pool).
- GPU0 ~93/97G, ~4.3G margin (gen fixed-util + static KV = stable, no OOM risk).
- New .env knobs: CHARRP_CTX / CHARRP_REASONING_CTX / CHARRP_KV_TYPE / CHARRP_REASONING_KV_TYPE.
This commit is contained in:
vh
2026-07-08 07:29:49 -07:00
parent b268f93035
commit f5706046b1
3 changed files with 29 additions and 4 deletions
+15 -2
View File
@@ -69,9 +69,15 @@ services:
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_CTX:-16384}"
- "${CHARRP_CTX:-98304}"
- --flash-attn
- on
# q8_0 KV cache ~halves KV VRAM (8-bit, near-lossless) → ~2x the context per GB.
# Mistral/Magistral handles q8 KV cleanly. Set f16 in .env to disable.
- --cache-type-k
- ${CHARRP_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_KV_TYPE:-q8_0}
- --jinja
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
@@ -115,9 +121,16 @@ services:
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_REASONING_CTX:-16384}"
- "${CHARRP_REASONING_CTX:-40960}"
- --flash-attn
- on
# QwQ native ctx = 40960 (beyond needs YaRN → quality loss; don't). Qwen-arch can be
# KV-quant-sensitive: q8_0 is verified coherent here, but flip to f16 in .env if a
# future model shows gibberish.
- --cache-type-k
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --jinja
# QwQ reasoning is template-native → llama.cpp manages it. --reasoning on surfaces
# the trace in reasoning_content (content stays clean prose); --reasoning-budget