feat(char-rp-reasoning): Deckard-PKD (Qwen3.5) replaces RpR-v4 after autonomous A/B
Operator wanted a reasoning-RP model that tolerates DRY (RpR-v4 forbids rep/DRY -> a 1/30 loop tail). Ran the full A/B on brokkr's 30-prompt D1 suite (content-only, slop-scored): - Deckard-PKD (Qwen3.5-27B, DavidAU creative tune) WON: 0/30 loops, 0/30 refusals, clean managed reasoning (native Qwen3.5 <think>/enable_thinking), DRY-tolerant, ~57 tok/s, runs on the base llama-swap b8840 image. -> now the char-rp-reasoning seat (:8018). - RpR-v4: 0 refusals but 1/30 loop (no-DRY). Pantheon-27B: clean slop but 7/30 explicit refusals + needs the newer ggml-org/llama.cpp image (Qwen3.6 won't load on b8840). Snowdrop + Gembrain (Gemma-4): floored (llama.cpp can't manage their reasoning without the vetoed template hacks). Losers kept on disk as alternates. - char-rp (Magidonia) unchanged; gen unchanged. gateway char-rp-reasoning -> Deckard sampler (temp 1.0/top_p 0.95/top_k 40/min_p 0.05; DRY server-side).
This commit is contained in:
@@ -35,21 +35,25 @@ CHARRP_KV_TYPE=q8_0
|
||||
# ── REASONING seat (char-rp-reasoning) ──────────────────────────────────────
|
||||
CHARRP_REASONING_CONTAINER=llama-charrp-reasoning
|
||||
CHARRP_REASONING_PORT=8018
|
||||
# Default = QwQ-32B-ArliAI-RpR-v4 Q5_K_M (abliterated reasoning → no re-censor;
|
||||
# llama.cpp-managed CoT). ~50 tok/s @ Q5. Use Q6_K (~46 tok/s) for a touch more
|
||||
# quality if speed is not binding.
|
||||
CHARRP_REASONING_MODEL=rp/QwQ-32B-ArliAI-RpR-v4-Q5_K_M.gguf
|
||||
# QwQ native ctx = 40960 (its max without YaRN). ~5.5G VRAM @ q8_0 KV.
|
||||
# Default = Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking i1-Q5_K_M (DavidAU creative
|
||||
# tune, native Qwen3.5 managed reasoning, DRY-tolerant). A/B WINNER 2026-07-08: 0/30 loops,
|
||||
# 0/30 refusals, clean slop; beat RpR-v4 (1/30 loop, no-DRY), Pantheon (7/30 refusals),
|
||||
# Snowdrop + Gembrain (template-incompatible with llama.cpp managed reasoning).
|
||||
CHARRP_REASONING_MODEL=rp/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking.i1-Q5_K_M.gguf
|
||||
# Qwen3.5-27B native ctx is large; 40960 = a sane reasoning-seat cap. ~5G VRAM @ q8_0 KV.
|
||||
CHARRP_REASONING_CTX=40960
|
||||
# KV cache dtype (Qwen-arch): q8_0 verified coherent here; f16 if a future model gibbers.
|
||||
# KV cache dtype: q8_0 verified coherent; f16 if a future model gibbers.
|
||||
CHARRP_REASONING_KV_TYPE=q8_0
|
||||
# Thinking-token cap (QwQ over-thinks otherwise → starves the prose). 300-500 = a
|
||||
# concise, useful scene-plan before the response.
|
||||
# Thinking-token cap (concise scene-plan before the response). 300-500 is a good band.
|
||||
CHARRP_REASONING_BUDGET=400
|
||||
# DRY anti-repetition multiplier (0 disables). Deckard tolerates DRY; DRY is the correct
|
||||
# anti-loop tool for the Qwen/QwQ family (a repetition PENALTY worsens their looping).
|
||||
CHARRP_REASONING_DRY=0.8
|
||||
|
||||
# ── ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning) ─────────
|
||||
# To collapse both seats onto Magidonia (drop QwQ): set
|
||||
# CHARRP_REASONING_MODEL=rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf
|
||||
# and remove the --reasoning* flags from the reasoning service in compose.yaml
|
||||
# (Magistral reasons only when the caller's system prompt contains "/think";
|
||||
# managed but LIGHT — see the compose header for why QwQ is the default).
|
||||
# ── ALTERNATES / FALLBACKS (all pre-pulled to /tank/aimodels/llm/rp/) ────────
|
||||
# Reasoning-seat A/B losers, kept on disk: QwQ-32B-ArliAI-RpR-v4-Q5_K_M (no-DRY → 1/30 loop
|
||||
# tail); Gryphe_Pantheon-Reasoning-27B-Q5_K_M (7/30 explicit refusals; also needs the newer
|
||||
# ggml-org/llama.cpp:server-cuda image — Qwen3.6 won't load on llama-swap b8840);
|
||||
# Gemma-4-Gembrain-31B (Gemma-4 reasoning-parser broken on llama.cpp — floored).
|
||||
# One-model fallback (collapse the reasoning seat onto Magidonia, lighter /think reasoning):
|
||||
# CHARRP_REASONING_MODEL=rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf + drop the DRY/--reasoning flags.
|
||||
|
||||
Reference in New Issue
Block a user