feat(char-rp-reasoning): Deckard-PKD (Qwen3.5) replaces RpR-v4 after autonomous A/B
Operator wanted a reasoning-RP model that tolerates DRY (RpR-v4 forbids rep/DRY -> a 1/30 loop tail). Ran the full A/B on brokkr's 30-prompt D1 suite (content-only, slop-scored): - Deckard-PKD (Qwen3.5-27B, DavidAU creative tune) WON: 0/30 loops, 0/30 refusals, clean managed reasoning (native Qwen3.5 <think>/enable_thinking), DRY-tolerant, ~57 tok/s, runs on the base llama-swap b8840 image. -> now the char-rp-reasoning seat (:8018). - RpR-v4: 0 refusals but 1/30 loop (no-DRY). Pantheon-27B: clean slop but 7/30 explicit refusals + needs the newer ggml-org/llama.cpp image (Qwen3.6 won't load on b8840). Snowdrop + Gembrain (Gemma-4): floored (llama.cpp can't manage their reasoning without the vetoed template hacks). Losers kept on disk as alternates. - char-rp (Magidonia) unchanged; gen unchanged. gateway char-rp-reasoning -> Deckard sampler (temp 1.0/top_p 0.95/top_k 40/min_p 0.05; DRY server-side).
This commit is contained in:
+17
-5
@@ -120,13 +120,15 @@ _As of 2026-07-08 — OFF-THE-SHELF INFERENCE STACK is the active work (home-tra
|
|||||||
elite literary prose, zero refusal on dark/explicit scenes, ~65 tok/s, tight POV/instruction adherence
|
elite literary prose, zero refusal on dark/explicit scenes, ~65 tok/s, tight POV/instruction adherence
|
||||||
(live-tested). Replaced the broken Angel NVFP4. Alt prose model (`.env` swap `CHARRP_MODEL`):
|
(live-tested). Replaced the broken Angel NVFP4. Alt prose model (`.env` swap `CHARRP_MODEL`):
|
||||||
`MS3.2-PaintedFantasy-v4.1-24B` (more literary flair, looser POV). Both GGUFs pre-pulled at `/tank/aimodels/llm/rp/`.
|
`MS3.2-PaintedFantasy-v4.1-24B` (more literary flair, looser POV). Both GGUFs pre-pulled at `/tank/aimodels/llm/rp/`.
|
||||||
- **char-rp-reasoning (:8018) = `QwQ-32B-ArliAI-RpR-v4-Q5_K_M` GGUF — LIVE.** QwQ reasoning RP tune (container
|
- **char-rp-reasoning (:8018) = `Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking.i1-Q5_K_M` GGUF — LIVE
|
||||||
`llama-charrp-reasoning`). `--reasoning on` → managed CoT SURFACES in `reasoning_content`, content stays clean
|
(A/B WINNER 2026-07-08, REPLACED RpR-v4).** DavidAU creative (PKD) tune on Qwen3.5-27B (container
|
||||||
prose, `--reasoning-budget 400` caps it. Reasoning data is QwQ-ABLITERATED → **no re-censor in `<think>`** (the
|
`llama-charrp-reasoning`). Native Qwen3.5 template → `--reasoning on` managed CoT SURFACES in `reasoning_content`,
|
||||||
Pantheon failure mode). ~52 tok/s @ Q5 (46 @ Q6). Best-of-breed-per-seat (NOT the same model as char-rp).
|
content stays clean prose, `--reasoning-budget 400` caps it; **DRY server-side** (sampler order = dry-after-temp)
|
||||||
|
tames looping. A/B: **0/30 loops + 0/30 refusals**, ~57 tok/s. Runs on the base llama-swap b8840 image (Qwen3.5).
|
||||||
|
Best-of-breed-per-seat (NOT the same model as char-rp).
|
||||||
|
|
||||||
- **🦄 RP-SEAT UNICORN — RESOLVED + DEPLOYED (2026-07-08).** char-rp = **Magidonia-24B-v4.3** (Magistral prose,
|
- **🦄 RP-SEAT UNICORN — RESOLVED + DEPLOYED (2026-07-08).** char-rp = **Magidonia-24B-v4.3** (Magistral prose,
|
||||||
65tps); char-rp-reasoning = **QwQ-32B-ArliAI-RpR-v4** (abliterated managed reasoning, 52tps @ Q5). Both GGUF via
|
65tps); char-rp-reasoning = **Deckard-PKD (Qwen3.5-27B)** (managed reasoning + DRY, ~57tps; A/B winner over RpR-v4 — see re-A/B note below). Both GGUF via
|
||||||
llama.cpp (`char-rp-gguf` stack, ana-ml2 GPU0, ~86/97G co-resident with gen, ~11G margin). Canonical stack in repo
|
llama.cpp (`char-rp-gguf` stack, ana-ml2 GPU0, ~86/97G co-resident with gen, ~11G margin). Canonical stack in repo
|
||||||
`stacks/char-rp-gguf/`; gateway rewired (`char-rp`→:8016, `char-rp-reasoning`→:8018, Mistral/QwQ samplers, dropped
|
`stacks/char-rp-gguf/`; gateway rewired (`char-rp`→:8016, `char-rp-reasoning`→:8018, Mistral/QwQ samplers, dropped
|
||||||
the Qwen `enable_thinking` kwarg). **KEY FINDINGS:** (a) no single dense 24-32B is BOTH an elite non-thinking prose
|
the Qwen `enable_thinking` kwarg). **KEY FINDINGS:** (a) no single dense 24-32B is BOTH an elite non-thinking prose
|
||||||
@@ -138,6 +140,16 @@ _As of 2026-07-08 — OFF-THE-SHELF INFERENCE STACK is the active work (home-tra
|
|||||||
old AEON trace-not-surfacing gap). One-model fallback (Magidonia both, lighter reasoning) documented in the stack
|
old AEON trace-not-surfacing gap). One-model fallback (Magidonia both, lighter reasoning) documented in the stack
|
||||||
header/README. Requirements met: prose#1, ≥50tps#2, low-refusal#3, dense#4, GGUF-not-Ollama#5, fit-GPU0#6, thinking#7.
|
header/README. Requirements met: prose#1, ≥50tps#2, low-refusal#3, dense#4, GGUF-not-Ollama#5, fit-GPU0#6, thinking#7.
|
||||||
Candidate GGUFs also on disk for A/B: Cydonia-R1-24B-v4.1, PaintedFantasy-v4.1-24B, RpR-v4 Q6_K.
|
Candidate GGUFs also on disk for A/B: Cydonia-R1-24B-v4.1, PaintedFantasy-v4.1-24B, RpR-v4 Q6_K.
|
||||||
|
**REASONING-SEAT RE-A/B (2026-07-08, later — RpR→Deckard):** operator wanted a reasoning model that TAKES DRY
|
||||||
|
(RpR forbids rep/DRY → 1/30 loop tail). Full A/B on brokkr's 30-prompt D1 suite (content-only, slop-scored):
|
||||||
|
**Deckard-PKD (Qwen3.5-27B) WON** (0/30 loops, 0/30 refusals, clean) → NOW the char-rp-reasoning seat.
|
||||||
|
RpR-v4: 0 refusals but 1/30 loop (no-DRY). **Pantheon-Reasoning-27B: 7/30 explicit refusals** (DeepSeek-distilled
|
||||||
|
re-censor IS real across the batch, milder than feared; NOT wholesale-rejected) — kept on disk as alt.
|
||||||
|
Snowdrop-v0.5-Type-S + Gembrain-31B (Gemma-4): FLOORED — llama.cpp can't do MANAGED reasoning on them (Snowdrop's
|
||||||
|
ChatML template has no `<think>`/`enable_thinking` hook; Gemma-4's reasoning parser splits content wrong). GATE for a
|
||||||
|
llama.cpp reasoning seat = STOCK template natively opens `<think>` or has `enable_thinking` (Qwen3.x/QwQ do; ChatML +
|
||||||
|
Gemma-4 don't). **INFRA:** llama-swap b8840 image can't load Qwen3.6/Gemma-4 archs → use
|
||||||
|
`ghcr.io/ggml-org/llama.cpp:server-cuda` (newer, pulled on ana-ml2) for those; Deckard (Qwen3.5) runs on b8840.
|
||||||
**MAX CONTEXT (2026-07-08):** char-rp **128K** (Magidonia FULL native 131072), char-rp-reasoning **40K** (QwQ
|
**MAX CONTEXT (2026-07-08):** char-rp **128K** (Magidonia FULL native 131072), char-rp-reasoning **40K** (QwQ
|
||||||
native 40960, YaRN-free max), **q8_0 KV cache both** (near-lossless, ~2× ctx/GB; verified coherent, no Qwen
|
native 40960, YaRN-free max), **q8_0 KV cache both** (near-lossless, ~2× ctx/GB; verified coherent, no Qwen
|
||||||
gibberish). **Funded by gen util 0.40→0.37** (freed ~2.9G of gen's IDLE KV headroom — gen KV usage runs 0-2%,
|
gibberish). **Funded by gen util 0.40→0.37** (freed ~2.9G of gen's IDLE KV headroom — gen KV usage runs 0-2%,
|
||||||
|
|||||||
@@ -35,21 +35,25 @@ CHARRP_KV_TYPE=q8_0
|
|||||||
# ── REASONING seat (char-rp-reasoning) ──────────────────────────────────────
|
# ── REASONING seat (char-rp-reasoning) ──────────────────────────────────────
|
||||||
CHARRP_REASONING_CONTAINER=llama-charrp-reasoning
|
CHARRP_REASONING_CONTAINER=llama-charrp-reasoning
|
||||||
CHARRP_REASONING_PORT=8018
|
CHARRP_REASONING_PORT=8018
|
||||||
# Default = QwQ-32B-ArliAI-RpR-v4 Q5_K_M (abliterated reasoning → no re-censor;
|
# Default = Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking i1-Q5_K_M (DavidAU creative
|
||||||
# llama.cpp-managed CoT). ~50 tok/s @ Q5. Use Q6_K (~46 tok/s) for a touch more
|
# tune, native Qwen3.5 managed reasoning, DRY-tolerant). A/B WINNER 2026-07-08: 0/30 loops,
|
||||||
# quality if speed is not binding.
|
# 0/30 refusals, clean slop; beat RpR-v4 (1/30 loop, no-DRY), Pantheon (7/30 refusals),
|
||||||
CHARRP_REASONING_MODEL=rp/QwQ-32B-ArliAI-RpR-v4-Q5_K_M.gguf
|
# Snowdrop + Gembrain (template-incompatible with llama.cpp managed reasoning).
|
||||||
# QwQ native ctx = 40960 (its max without YaRN). ~5.5G VRAM @ q8_0 KV.
|
CHARRP_REASONING_MODEL=rp/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking.i1-Q5_K_M.gguf
|
||||||
|
# Qwen3.5-27B native ctx is large; 40960 = a sane reasoning-seat cap. ~5G VRAM @ q8_0 KV.
|
||||||
CHARRP_REASONING_CTX=40960
|
CHARRP_REASONING_CTX=40960
|
||||||
# KV cache dtype (Qwen-arch): q8_0 verified coherent here; f16 if a future model gibbers.
|
# KV cache dtype: q8_0 verified coherent; f16 if a future model gibbers.
|
||||||
CHARRP_REASONING_KV_TYPE=q8_0
|
CHARRP_REASONING_KV_TYPE=q8_0
|
||||||
# Thinking-token cap (QwQ over-thinks otherwise → starves the prose). 300-500 = a
|
# Thinking-token cap (concise scene-plan before the response). 300-500 is a good band.
|
||||||
# concise, useful scene-plan before the response.
|
|
||||||
CHARRP_REASONING_BUDGET=400
|
CHARRP_REASONING_BUDGET=400
|
||||||
|
# DRY anti-repetition multiplier (0 disables). Deckard tolerates DRY; DRY is the correct
|
||||||
|
# anti-loop tool for the Qwen/QwQ family (a repetition PENALTY worsens their looping).
|
||||||
|
CHARRP_REASONING_DRY=0.8
|
||||||
|
|
||||||
# ── ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning) ─────────
|
# ── ALTERNATES / FALLBACKS (all pre-pulled to /tank/aimodels/llm/rp/) ────────
|
||||||
# To collapse both seats onto Magidonia (drop QwQ): set
|
# Reasoning-seat A/B losers, kept on disk: QwQ-32B-ArliAI-RpR-v4-Q5_K_M (no-DRY → 1/30 loop
|
||||||
# CHARRP_REASONING_MODEL=rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf
|
# tail); Gryphe_Pantheon-Reasoning-27B-Q5_K_M (7/30 explicit refusals; also needs the newer
|
||||||
# and remove the --reasoning* flags from the reasoning service in compose.yaml
|
# ggml-org/llama.cpp:server-cuda image — Qwen3.6 won't load on llama-swap b8840);
|
||||||
# (Magistral reasons only when the caller's system prompt contains "/think";
|
# Gemma-4-Gembrain-31B (Gemma-4 reasoning-parser broken on llama.cpp — floored).
|
||||||
# managed but LIGHT — see the compose header for why QwQ is the default).
|
# One-model fallback (collapse the reasoning seat onto Magidonia, lighter /think reasoning):
|
||||||
|
# CHARRP_REASONING_MODEL=rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf + drop the DRY/--reasoning flags.
|
||||||
|
|||||||
@@ -94,7 +94,7 @@ services:
|
|||||||
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
|
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
|
||||||
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
|
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
|
||||||
|
|
||||||
# ── REASONING seat — QwQ managed thinking. gateway char-rp-reasoning. ──
|
# ── REASONING seat — Deckard-PKD (Qwen3.5) managed thinking + DRY. gateway char-rp-reasoning. ──
|
||||||
llama-charrp-reasoning:
|
llama-charrp-reasoning:
|
||||||
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
||||||
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
|
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
|
||||||
@@ -113,7 +113,7 @@ services:
|
|||||||
entrypoint: ["/app/llama-server"]
|
entrypoint: ["/app/llama-server"]
|
||||||
command:
|
command:
|
||||||
- --model
|
- --model
|
||||||
- /models/${CHARRP_REASONING_MODEL:-rp/QwQ-32B-ArliAI-RpR-v4-Q5_K_M.gguf}
|
- /models/${CHARRP_REASONING_MODEL:-rp/Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking.i1-Q5_K_M.gguf}
|
||||||
- --host
|
- --host
|
||||||
- 0.0.0.0
|
- 0.0.0.0
|
||||||
- --port
|
- --port
|
||||||
@@ -124,23 +124,33 @@ services:
|
|||||||
- "${CHARRP_REASONING_CTX:-40960}"
|
- "${CHARRP_REASONING_CTX:-40960}"
|
||||||
- --flash-attn
|
- --flash-attn
|
||||||
- on
|
- on
|
||||||
# QwQ native ctx = 40960 (beyond needs YaRN → quality loss; don't). Qwen-arch can be
|
# Deckard = Qwen3.5-27B (native ctx large); 40960 is a sane reasoning-seat cap. q8_0 KV
|
||||||
# KV-quant-sensitive: q8_0 is verified coherent here, but flip to f16 in .env if a
|
# verified coherent; flip to f16 in .env if a future model shows gibberish.
|
||||||
# future model shows gibberish.
|
|
||||||
- --cache-type-k
|
- --cache-type-k
|
||||||
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
||||||
- --cache-type-v
|
- --cache-type-v
|
||||||
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
|
||||||
- --jinja
|
- --jinja
|
||||||
# QwQ reasoning is template-native → llama.cpp manages it. --reasoning on surfaces
|
# Deckard's Qwen3.5 template natively opens <think> + has enable_thinking → llama.cpp
|
||||||
# the trace in reasoning_content (content stays clean prose); --reasoning-budget
|
# manages the reasoning (trace to reasoning_content, content stays clean prose);
|
||||||
# caps the CoT so it can't run away and starve the prose (QwQ over-thinks otherwise).
|
# --reasoning-budget caps the CoT. (A/B 2026-07-08: Deckard 0/30 loops + 0/30 refusals,
|
||||||
|
# beat RpR-v4 (1/30 loop), Pantheon (7/30 refusals), Snowdrop + Gembrain (template-incompat).)
|
||||||
- --reasoning
|
- --reasoning
|
||||||
- on
|
- on
|
||||||
- --reasoning-format
|
- --reasoning-format
|
||||||
- deepseek
|
- deepseek
|
||||||
- --reasoning-budget
|
- --reasoning-budget
|
||||||
- "${CHARRP_REASONING_BUDGET:-400}"
|
- "${CHARRP_REASONING_BUDGET:-400}"
|
||||||
|
# DRY anti-repetition (Deckard/Qwen3.5 tolerates it; sampler ORDER = dry AFTER temperature,
|
||||||
|
# per the QwQ-family finding that a repetition PENALTY worsens looping but DRY fixes it).
|
||||||
|
- --samplers
|
||||||
|
- "top_k;top_p;min_p;temperature;dry"
|
||||||
|
- --dry-multiplier
|
||||||
|
- "${CHARRP_REASONING_DRY:-0.8}"
|
||||||
|
- --dry-base
|
||||||
|
- "1.75"
|
||||||
|
- --dry-allowed-length
|
||||||
|
- "2"
|
||||||
healthcheck:
|
healthcheck:
|
||||||
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
||||||
interval: 30s
|
interval: 30s
|
||||||
|
|||||||
@@ -173,23 +173,23 @@ model_list:
|
|||||||
model_info:
|
model_info:
|
||||||
mode: chat
|
mode: chat
|
||||||
# char-rp-reasoning -> GGUF managed-REASONING seat (:8018, llama.cpp, char-rp-gguf stack).
|
# char-rp-reasoning -> GGUF managed-REASONING seat (:8018, llama.cpp, char-rp-gguf stack).
|
||||||
# ArliAI QwQ-32B-ArliAI-RpR-v4 Q5_K_M — QwQ reasoning RP tune. Reasoning is ON server-side
|
# Qwen3.5-27B-Deckard-PKD-Heretic-Uncensored-Thinking i1-Q5_K_M — DavidAU creative tune.
|
||||||
# (--reasoning on): the CoT SURFACES in reasoning_content and content stays clean prose
|
# Reasoning ON server-side (--reasoning on): CoT surfaces in reasoning_content, content stays
|
||||||
# (fixes the old trace-not-surfacing gap), CoT budget-capped so it can't starve the prose.
|
# clean prose, budget-capped. DRY server-side (sampler order = dry after temperature) tames looping.
|
||||||
# Its reasoning data is QwQ-ABLITERATED → no re-censor inside <think> (the failure mode
|
# A/B WINNER 2026-07-08: 0/30 loops + 0/30 refusals; beat RpR-v4 (1/30 loop, forbids DRY),
|
||||||
# that disqualified Pantheon-Reasoning-27B). ~52 tok/s @ Q5_K_M. RpR card: temp 1.0,
|
# Pantheon-27B (7/30 explicit refusals), Snowdrop + Gembrain (llama.cpp template-incompat).
|
||||||
# top_k 40, min_p 0.02, and NO repetition / DRY / XTC penalties. NOT the same model as
|
# Deckard decode: temp 1.0, top_p 0.95, top_k 40, min_p 0.05. NOT the same model as char-rp
|
||||||
# char-rp (best-of-breed per seat) — see stacks/char-rp-gguf/README.md.
|
# (best-of-breed per seat) — see stacks/char-rp-gguf/README.md.
|
||||||
- model_name: char-rp-reasoning
|
- model_name: char-rp-reasoning
|
||||||
litellm_params:
|
litellm_params:
|
||||||
model: hosted_vllm/qwq-32b-rpr-v4
|
model: hosted_vllm/deckard-pkd-27b
|
||||||
api_base: http://10.250.50.54:8018/v1
|
api_base: http://10.250.50.54:8018/v1
|
||||||
api_key: os.environ/VLLM_API_KEY
|
api_key: os.environ/VLLM_API_KEY
|
||||||
temperature: 1.0
|
temperature: 1.0
|
||||||
top_p: 0.95
|
top_p: 0.95
|
||||||
extra_body:
|
extra_body:
|
||||||
top_k: 40
|
top_k: 40
|
||||||
min_p: 0.02
|
min_p: 0.05
|
||||||
model_info:
|
model_info:
|
||||||
mode: chat
|
mode: chat
|
||||||
|
|
||||||
|
|||||||
Reference in New Issue
Block a user