feat(char-rp-reasoning): Deckard-PKD (Qwen3.5) replaces RpR-v4 after autonomous A/B
Operator wanted a reasoning-RP model that tolerates DRY (RpR-v4 forbids rep/DRY -> a 1/30 loop tail). Ran the full A/B on brokkr's 30-prompt D1 suite (content-only, slop-scored): - Deckard-PKD (Qwen3.5-27B, DavidAU creative tune) WON: 0/30 loops, 0/30 refusals, clean managed reasoning (native Qwen3.5 <think>/enable_thinking), DRY-tolerant, ~57 tok/s, runs on the base llama-swap b8840 image. -> now the char-rp-reasoning seat (:8018). - RpR-v4: 0 refusals but 1/30 loop (no-DRY). Pantheon-27B: clean slop but 7/30 explicit refusals + needs the newer ggml-org/llama.cpp image (Qwen3.6 won't load on b8840). Snowdrop + Gembrain (Gemma-4): floored (llama.cpp can't manage their reasoning without the vetoed template hacks). Losers kept on disk as alternates. - char-rp (Magidonia) unchanged; gen unchanged. gateway char-rp-reasoning -> Deckard sampler (temp 1.0/top_p 0.95/top_k 40/min_p 0.05; DRY server-side).
This commit is contained in:
@@ -94,7 +94,7 @@ services:
|
||||
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
|
||||
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
|
||||
|
||||
# ── REASONING seat — QwQ managed thinking. gateway char-rp-reasoning. ──
|
||||
# ── REASONING seat — Deckard-PKD (Qwen3.5) managed thinking + DRY. gateway char-rp-reasoning. ──
|
||||
llama-charrp-reasoning:
|
||||
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
||||
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
|
||||
@@ -113,7 +113,7 @@ services:
|
||||
entrypoint: ["/app/llama-server"]
|
||||
command:
|
||||
- --model
|
||||
- /models/${CHARRP_REASONING_MODEL:-rp/QwQ-32B-ArliAI-RpR-v4-Q5_K_M.gguf}
|
||||
- /models/${CHARRP_REASONING_MODEL:-rp/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking.i1-Q5_K_M.gguf}
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
@@ -124,23 +124,33 @@ services:
|
||||
- "${CHARRP_REASONING_CTX:-40960}"
|
||||
- --flash-attn
|
||||
- on
|
||||
# QwQ native ctx = 40960 (beyond needs YaRN → quality loss; don't). Qwen-arch can be
|
||||
# KV-quant-sensitive: q8_0 is verified coherent here, but flip to f16 in .env if a
|
||||
# future model shows gibberish.
|
||||
# Deckard = Qwen3.5-27B (native ctx large); 40960 is a sane reasoning-seat cap. q8_0 KV
|
||||
# verified coherent; flip to f16 in .env if a future model shows gibberish.
|
||||
- --cache-type-k
|
||||
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
||||
- --cache-type-v
|
||||
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
||||
- --jinja
|
||||
# QwQ reasoning is template-native → llama.cpp manages it. --reasoning on surfaces
|
||||
# the trace in reasoning_content (content stays clean prose); --reasoning-budget
|
||||
# caps the CoT so it can't run away and starve the prose (QwQ over-thinks otherwise).
|
||||
# Deckard's Qwen3.5 template natively opens <think> + has enable_thinking → llama.cpp
|
||||
# manages the reasoning (trace to reasoning_content, content stays clean prose);
|
||||
# --reasoning-budget caps the CoT. (A/B 2026-07-08: Deckard 0/30 loops + 0/30 refusals,
|
||||
# beat RpR-v4 (1/30 loop), Pantheon (7/30 refusals), Snowdrop + Gembrain (template-incompat).)
|
||||
- --reasoning
|
||||
- on
|
||||
- --reasoning-format
|
||||
- deepseek
|
||||
- --reasoning-budget
|
||||
- "${CHARRP_REASONING_BUDGET:-400}"
|
||||
# DRY anti-repetition (Deckard/Qwen3.5 tolerates it; sampler ORDER = dry AFTER temperature,
|
||||
# per the QwQ-family finding that a repetition PENALTY worsens looping but DRY fixes it).
|
||||
- --samplers
|
||||
- "top_k;top_p;min_p;temperature;dry"
|
||||
- --dry-multiplier
|
||||
- "${CHARRP_REASONING_DRY:-0.8}"
|
||||
- --dry-base
|
||||
- "1.75"
|
||||
- --dry-allowed-length
|
||||
- "2"
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
||||
interval: 30s
|
||||
|
||||
Reference in New Issue
Block a user