feat(char-rp-gguf): replace broken Angel NVFP4 with dual GGUF RP seat on ana-ml2 GPU0
char-rp -> TheDrummer Magidonia-24B-v4.3 Q6_K (Magistral prose, ~65 tok/s,
zero refusal, tight POV) via llama.cpp (:8016).
char-rp-reasoning -> ArliAI QwQ-32B-RpR-v4 Q5_K_M (abliterated managed reasoning,
~52 tok/s, reasoning surfaces in reasoning_content) via llama.cpp (:8018).
- New canonical stack stacks/char-rp-gguf/ (llama-server x2, GPU0-pinned, ~86/97G
co-resident with gen). GGUF sidesteps the vLLM-NVFP4 + Mistral-tokenizer traps that
killed the Angel serve. Never Ollama.
- Best-of-breed per seat: no single dense 24-32B is both an elite non-thinking prose
seat AND a clean managed-reasoning seat on llama.cpp (Magidonia [THINK] boundary is
loose; Cydonia-R1 <think> runs away; QwQ is template-managed). Pantheon-Reasoning-27B
stays rejected (re-censors in <think>; RpR-v4 abliterated reasoning is the fix).
- Gateway rewired: char-rp->:8016, char-rp-reasoning->:8018, Mistral/QwQ samplers,
dropped the Qwen enable_thinking kwarg. One-model Magidonia fallback documented.
- Retired the ms32-24b-angel stack.
This commit is contained in:
@@ -0,0 +1,149 @@
|
||||
# char-rp-gguf — dedicated GGUF character-RP seat on ana-ml2 GPU 0, REPLACING the
|
||||
# broken ms32-24b-angel NVFP4 serve (garbage output — bad self-quant W4A4).
|
||||
#
|
||||
# Two co-located llama.cpp (llama-server) instances on GPU 0, served alongside the
|
||||
# 35B-A3B heretic `gen` (qwen36-27b-aeon stack, :8015):
|
||||
#
|
||||
# llama-charrp (:8016, gateway char-rp) — TheDrummer Magidonia-24B-v4.3 Q6_K.
|
||||
# Magistral (Mistral) dark-romantasy RP tune. NON-thinking PROSE seat: elite
|
||||
# literary prose, zero refusal, ~65 tok/s, precise POV/instruction adherence.
|
||||
#
|
||||
# llama-charrp-reasoning (:8018, gateway char-rp-reasoning) — ArliAI QwQ-32B-RpR-v4 Q5_K_M.
|
||||
# QwQ reasoning RP tune whose reasoning DATA was generated with QwQ-ABLITERATED
|
||||
# → it does NOT re-censor in the think phase (the exact failure mode that killed
|
||||
# the Pantheon/DeepSeek-distilled reasoners: they reason themselves into refusals
|
||||
# inside <think>). llama.cpp MANAGES QwQ reasoning natively: --reasoning on
|
||||
# surfaces the trace in reasoning_content (clean prose in content, no <think>
|
||||
# leak), --reasoning-budget caps the chain-of-thought. ~50 tok/s @ Q5_K_M.
|
||||
#
|
||||
# WHY GGUF/llama.cpp (not vLLM NVFP4): sidesteps BOTH traps that killed the Angel serve
|
||||
# — the vLLM NVFP4 self-quant breakage AND the Mistral-tokenizer/vision crash. llama.cpp
|
||||
# handles Mistral + QwQ tokenizers natively. NEVER Ollama (banned fleet-wide).
|
||||
#
|
||||
# WHY TWO models (not one): no single dense 24-32B is BOTH an elite non-thinking prose
|
||||
# seat AND a clean managed-reasoning seat on llama.cpp. Magidonia's Magistral [THINK]
|
||||
# discipline is loose (won't reliably close [/THINK] on substantive reasoning → prose
|
||||
# bleeds into reasoning_content, content empties); Cydonia-R1's <think> is emergent, so
|
||||
# llama.cpp can't manage/cap it → runaway CoT that never reaches prose. QwQ's template
|
||||
# opens <think> natively → llama.cpp manages+caps it. So: best-of-breed per seat.
|
||||
# ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning): point both services
|
||||
# at Magidonia via CHARRP_REASONING_MODEL in .env and blank CHARRP_REASONING_EXTRA_*.
|
||||
#
|
||||
# ALTERNATE prose model: PaintedFantasy-v4.1-24B (also Magistral, more literary flair
|
||||
# but looser POV adherence) — set CHARRP_MODEL in .env. All candidate GGUFs are
|
||||
# pre-pulled to /tank/aimodels/llm/rp/.
|
||||
#
|
||||
# VRAM (GPU 0, co-resident with gen ~38G): Magidonia Q6 ~19G + RpR-v4 Q5 ~23G + KV/
|
||||
# compute ~6-8G = ~85-88G / 97G (~9-12G margin). Keep ctx modest; drop CHARRP_*_CTX
|
||||
# to 8192 in .env if warmup bites. depends_on sequences char-rp first.
|
||||
#
|
||||
# API auth: blank (LAN-internal on the GPU host; matches API_KEY= in the AEON stack /
|
||||
# gateway VLLM_API_KEY). llama-server ignores the gateway's api_key when none is set.
|
||||
#
|
||||
# All tunables live in .env — edit that, not this file.
|
||||
|
||||
name: char-rp-gguf
|
||||
|
||||
services:
|
||||
# ── PROSE seat — non-thinking. gateway char-rp. ──
|
||||
llama-charrp:
|
||||
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
||||
container_name: ${CHARRP_CONTAINER:-llama-charrp}
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
ports:
|
||||
- "${CHARRP_PORT:-8016}:8080"
|
||||
volumes:
|
||||
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
|
||||
environment:
|
||||
# Pin to GPU 0 (the on-demand large-model card; the always-on vLLM trio owns GPU 1).
|
||||
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
|
||||
entrypoint: ["/app/llama-server"]
|
||||
command:
|
||||
- --model
|
||||
- /models/${CHARRP_MODEL:-rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf}
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8080"
|
||||
- --n-gpu-layers
|
||||
- "999"
|
||||
- --ctx-size
|
||||
- "${CHARRP_CTX:-16384}"
|
||||
- --flash-attn
|
||||
- on
|
||||
- --jinja
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 240s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=char-rp (Magidonia-24B GGUF)
|
||||
- homepage.icon=mdi-drama-masks
|
||||
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, ana-ml2 GPU 0)
|
||||
- homepage.href=http://10.250.50.54:${CHARRP_PORT:-8016}
|
||||
|
||||
# ── REASONING seat — QwQ managed thinking. gateway char-rp-reasoning. ──
|
||||
llama-charrp-reasoning:
|
||||
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
|
||||
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
|
||||
restart: unless-stopped
|
||||
runtime: nvidia
|
||||
# Sequence AFTER the prose seat is healthy so the two GPU-0 allocations don't race.
|
||||
depends_on:
|
||||
llama-charrp:
|
||||
condition: service_healthy
|
||||
ports:
|
||||
- "${CHARRP_REASONING_PORT:-8018}:8080"
|
||||
volumes:
|
||||
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
|
||||
environment:
|
||||
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
|
||||
entrypoint: ["/app/llama-server"]
|
||||
command:
|
||||
- --model
|
||||
- /models/${CHARRP_REASONING_MODEL:-rp/QwQ-32B-ArliAI-RpR-v4-Q5_K_M.gguf}
|
||||
- --host
|
||||
- 0.0.0.0
|
||||
- --port
|
||||
- "8080"
|
||||
- --n-gpu-layers
|
||||
- "999"
|
||||
- --ctx-size
|
||||
- "${CHARRP_REASONING_CTX:-16384}"
|
||||
- --flash-attn
|
||||
- on
|
||||
- --jinja
|
||||
# QwQ reasoning is template-native → llama.cpp manages it. --reasoning on surfaces
|
||||
# the trace in reasoning_content (content stays clean prose); --reasoning-budget
|
||||
# caps the CoT so it can't run away and starve the prose (QwQ over-thinks otherwise).
|
||||
- --reasoning
|
||||
- on
|
||||
- --reasoning-format
|
||||
- deepseek
|
||||
- --reasoning-budget
|
||||
- "${CHARRP_REASONING_BUDGET:-400}"
|
||||
healthcheck:
|
||||
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
|
||||
interval: 30s
|
||||
timeout: 10s
|
||||
retries: 3
|
||||
start_period: 300s
|
||||
networks:
|
||||
- tnet
|
||||
labels:
|
||||
- homepage.group=AI Systems
|
||||
- homepage.name=char-rp-reasoning (QwQ-32B RpR-v4 GGUF)
|
||||
- homepage.icon=mdi-brain
|
||||
- homepage.description=Dark-romantasy RP reasoning seat, managed CoT (llama.cpp, ana-ml2 GPU 0)
|
||||
- homepage.href=http://10.250.50.54:${CHARRP_REASONING_PORT:-8018}
|
||||
|
||||
networks:
|
||||
tnet:
|
||||
name: traefik-net
|
||||
external: true
|
||||
Reference in New Issue
Block a user