feat(char-rp-gguf): replace broken Angel NVFP4 with dual GGUF RP seat on ana-ml2 GPU0

char-rp        -> TheDrummer Magidonia-24B-v4.3 Q6_K (Magistral prose, ~65 tok/s,
                  zero refusal, tight POV) via llama.cpp (:8016).
char-rp-reasoning -> ArliAI QwQ-32B-RpR-v4 Q5_K_M (abliterated managed reasoning,
                  ~52 tok/s, reasoning surfaces in reasoning_content) via llama.cpp (:8018).

- New canonical stack stacks/char-rp-gguf/ (llama-server x2, GPU0-pinned, ~86/97G
  co-resident with gen). GGUF sidesteps the vLLM-NVFP4 + Mistral-tokenizer traps that
  killed the Angel serve. Never Ollama.
- Best-of-breed per seat: no single dense 24-32B is both an elite non-thinking prose
  seat AND a clean managed-reasoning seat on llama.cpp (Magidonia [THINK] boundary is
  loose; Cydonia-R1 <think> runs away; QwQ is template-managed). Pantheon-Reasoning-27B
  stays rejected (re-censors in <think>; RpR-v4 abliterated reasoning is the fix).
- Gateway rewired: char-rp->:8016, char-rp-reasoning->:8018, Mistral/QwQ samplers,
  dropped the Qwen enable_thinking kwarg. One-model Magidonia fallback documented.
- Retired the ms32-24b-angel stack.
This commit is contained in:
vh
2026-07-08 02:43:00 -07:00
parent 75851c2837
commit b268f93035
5 changed files with 500 additions and 55 deletions
+95
View File
@@ -0,0 +1,95 @@
# char-rp-gguf
Dedicated GGUF **character-RP seat** on ana-ml2 GPU 0, served via llama.cpp
(`llama-server`) — **never Ollama**. Replaces the broken `ms32-24b-angel` NVFP4
serve (self-quant produced garbage even at greedy).
Two co-located instances, alongside the 35B-A3B heretic `gen` (`qwen36-27b-aeon`,
:8015):
| service | port | gateway alias | model | role |
|---|---|---|---|---|
| `llama-charrp` | 8016 | `char-rp` | TheDrummer **Magidonia-24B-v4.3** Q6_K | non-thinking **prose** seat |
| `llama-charrp-reasoning` | 8018 | `char-rp-reasoning` | ArliAI **QwQ-32B-RpR-v4** Q5_K_M | managed-**reasoning** seat |
## Why these two models
Requirements (operator, 2026-07-08): highest prose quality; ≥50 tok/s (single-seat);
very low refusal for dark explicit romantasy; dense; GGUF-via-llama.cpp; thinking
strongly preferred; must fit alongside the 35B-A3B on GPU 0.
- **Magidonia-24B-v4.3** (`char-rp`) — Magistral (Mistral) dark-romantasy RP tune.
Live-tested: elite literary prose, **zero refusal** on explicit scenes, precise
POV/instruction adherence, **~65 tok/s** at Q6_K. The prose star.
- **QwQ-32B-RpR-v4** (`char-rp-reasoning`) — QwQ reasoning RP tune whose reasoning
data was generated with **QwQ-abliterated**, so it does **not re-censor inside the
think phase** — the failure mode that disqualified Pantheon-Reasoning-27B (its
DeepSeek-distilled traces reason themselves into refusals). llama.cpp **manages**
QwQ reasoning: `--reasoning on` surfaces the trace in `reasoning_content` (content
stays clean prose, no `<think>` leak) and `--reasoning-budget` caps the CoT.
~50 tok/s at Q5_K_M (46 at Q6).
**Why not one model for both:** no single dense 24–32B is both an elite non-thinking
prose seat and a clean *managed*-reasoning seat on llama.cpp. Magidonia's Magistral
`[THINK]` discipline is loose (won't reliably close `[/THINK]` on substantive
reasoning → prose bleeds into `reasoning_content`, `content` empties). Cydonia-R1's
`<think>` is emergent, so llama.cpp can't manage/cap it → runaway CoT that never
reaches the prose. QwQ's template opens `<think>` natively → llama.cpp manages+caps
it. Best-of-breed per seat. See the compose header for the one-model fallback.
**Alternate prose model:** `MS3.2-PaintedFantasy-v4.1-24B` (Magistral; more literary
flair, looser POV adherence) — swap via `CHARRP_MODEL` in `.env`. All candidate GGUFs
are pre-pulled to `/tank/aimodels/llm/rp/`.
## Deploy
Canonical copies live here; authoritative copies live on the host under
`/opt/docker/compose/char-rp-gguf/`.
```bash
# On ana-ml2 (models already pre-pulled to /tank/aimodels/llm/rp/):
scripts/deploy-stack.sh ana-ml2 char-rp-gguf # from this workspace
# or, on the host:
cd /opt/docker/compose/char-rp-gguf && cp .env.example .env && docker compose up -d
```
GGUFs are pulled with a detached `huggingface_hub` container, e.g.:
```bash
docker run -d --entrypoint python3 -v /tank:/tank vllm/vllm-openai:latest -c \
'from huggingface_hub import hf_hub_download as d; \
d(repo_id="bartowski/TheDrummer_Magidonia-24B-v4.3-GGUF", \
filename="TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf", local_dir="/tank/aimodels/llm/rp")'
```
## Gateway wiring (ana-docker LiteLLM `:4000`)
`char-rp` → `http://10.250.50.54:8016/v1`, `char-rp-reasoning` →
`http://10.250.50.54:8018/v1`. Sampler defaults (per the model cards):
- **char-rp (Magidonia / Mistral):** `temperature 1.0, top_p 0.95, min_p 0.03`;
no repetition penalty (Magistral is stable without one). Callers may override.
- **char-rp-reasoning (QwQ RpR-v4):** `temperature 1.0, top_k 40, min_p 0.02`;
**no repetition / DRY / XTC penalties** (RpR card is explicit about this).
Reasoning surfaces in `reasoning_content`.
Do **not** carry the Qwen-specific `chat_template_kwargs.enable_thinking` on these
aliases — it is meaningless to Mistral/QwQ templates. Reasoning is controlled
server-side (`--reasoning on` on :8018).
## Revert
```bash
docker compose -f /opt/docker/compose/char-rp-gguf/compose.yaml down
```
The retired `ms32-24b-angel` stack is left staged for reference (its NVFP4 quant is
broken — do not re-serve it).
## Verify
```bash
curl -s http://10.250.50.54:8016/v1/models # char-rp up
curl -s http://10.250.50.54:8018/v1/models # char-rp-reasoning up
nvidia-smi --query-gpu=index,memory.free --format=csv # GPU 0 margin
```