Files
vh 3132a16ca0 fv-ml1: finish the renumber the cutover missed -- 16 dead dashboard links
Every fv-ml1 link on the Homepage dashboard was broken. Measured against the
live dashboard API before the fix: 16 entries pointing at the dead 10.250.50.54
and zero at the live 10.251.50.54, covering gen, M.O.G.-SEC, Scriberr, Embed,
Rerank, Reward, Coder, Dockge and six dormant seats.

The miss was structural, not careless. fv-ml1-rename-sweep.sh works from an
allowlist assembled from files that mention the HOST, and a homepage.href label
mentions only an IP -- so every stack whose sole stale reference was a label
fell outside it. The allowlist now covers those 24 files, and records how to
derive the list next time (grep the old address, subtract history) rather than
enumerating from memory.

History is still untouched, and the exclusions are now written down with the
reason each one keeps the old address: recorded benchmark results, whose
base_url is part of a measurement's provenance; the one LiteLLM comment
preserving a retired hand-test endpoint; and the cutover runbooks, where the old
address is the subject matter.

Two bugs found while applying it, both fixed here:

  - deploy-stack.sh rejected any stack name containing a dot, so qwen3.5-122b,
    qwopus3.5-122b and mistral-medium-3.5 could not be deployed by the script at
    all. The check exists to stop path traversal, which means rejecting ".." and
    "/" -- not every dot. Traversal is now rejected explicitly and tested.
  - stacks/scriberr/.env.example allowed CORS only from the dead IP and from
    scriberr.ana.internal, which no longer resolves; the box is at the fv site
    and DNS already carries scriberr.fv.internal. The live .env had both stale
    origins, i.e. an allowlist with nothing reachable in it.

Host side, applied separately: canonical pushed for the 16 stacks whose only
difference from the host was this renumber, and an in-place address-only fix for
the nine whose host copy has genuinely drifted or has no canonical copy, so that
drift survives for a deliberate reconciliation instead of being clobbered. Every
compose.yaml on fv-ml1 now reads 10.251.50.54. The labels themselves only take
effect at container creation, so the running containers still need recreating.
2026-09-12 23:05:29 -07:00

182 lines
8.1 KiB
YAML

# char-rp-gguf — dedicated GGUF character-RP seat on fv-ml1 GPU 0, REPLACING the
# broken ms32-24b-angel NVFP4 serve (garbage output — bad self-quant W4A4).
#
# Two co-located llama.cpp (llama-server) instances on GPU 0, served alongside the
# 35B-A3B heretic `gen` (qwen36-27b-aeon stack, :8015):
#
# llama-charrp (:8016, gateway char-rp) — TheDrummer Magidonia-24B-v4.3 Q6_K.
# Magistral (Mistral) dark-romantasy RP tune. NON-thinking PROSE seat: elite
# literary prose, zero refusal, ~65 tok/s, precise POV/instruction adherence.
#
# llama-charrp-reasoning (:8018, gateway char-rp-reasoning) — ArliAI QwQ-32B-RpR-v4 Q5_K_M.
# QwQ reasoning RP tune whose reasoning DATA was generated with QwQ-ABLITERATED
# → it does NOT re-censor in the think phase (the exact failure mode that killed
# the Pantheon/DeepSeek-distilled reasoners: they reason themselves into refusals
# inside <think>). llama.cpp MANAGES QwQ reasoning natively: --reasoning on
# surfaces the trace in reasoning_content (clean prose in content, no <think>
# leak), --reasoning-budget caps the chain-of-thought. ~50 tok/s @ Q5_K_M.
#
# WHY GGUF/llama.cpp (not vLLM NVFP4): sidesteps BOTH traps that killed the Angel serve
# — the vLLM NVFP4 self-quant breakage AND the Mistral-tokenizer/vision crash. llama.cpp
# handles Mistral + QwQ tokenizers natively. NEVER Ollama (banned fleet-wide).
#
# WHY TWO models (not one): no single dense 24-32B is BOTH an elite non-thinking prose
# seat AND a clean managed-reasoning seat on llama.cpp. Magidonia's Magistral [THINK]
# discipline is loose (won't reliably close [/THINK] on substantive reasoning → prose
# bleeds into reasoning_content, content empties); Cydonia-R1's <think> is emergent, so
# llama.cpp can't manage/cap it → runaway CoT that never reaches prose. QwQ's template
# opens <think> natively → llama.cpp manages+caps it. So: best-of-breed per seat.
# ONE-MODEL FALLBACK (consistent Mistral style, lighter reasoning): point both services
# at Magidonia via CHARRP_REASONING_MODEL in .env and blank CHARRP_REASONING_EXTRA_*.
#
# ALTERNATE prose model: PaintedFantasy-v4.1-24B (also Magistral, more literary flair
# but looser POV adherence) — set CHARRP_MODEL in .env. All candidate GGUFs are
# pre-pulled to /tank/aimodels/llm/rp/.
#
# VRAM (GPU 0, co-resident with gen ~38G): Magidonia Q6 ~19G + RpR-v4 Q5 ~23G + KV/
# compute ~6-8G = ~85-88G / 97G (~9-12G margin). Keep ctx modest; drop CHARRP_*_CTX
# to 8192 in .env if warmup bites. depends_on sequences char-rp first.
#
# API auth: blank (LAN-internal on the GPU host; matches API_KEY= in the AEON stack /
# gateway VLLM_API_KEY). llama-server ignores the gateway's api_key when none is set.
#
# All tunables live in .env — edit that, not this file.
name: char-rp-gguf
services:
# ── PROSE seat — non-thinking. gateway char-rp. ──
llama-charrp:
image: ${LLAMA_IMAGE:-ghcr.io/mostlygeek/llama-swap:cuda}
container_name: ${CHARRP_CONTAINER:-llama-charrp}
restart: unless-stopped
runtime: nvidia
ports:
- "${CHARRP_PORT:-8016}:8080"
volumes:
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
environment:
# Pin to GPU 0 (the on-demand large-model card; the always-on vLLM trio owns GPU 1).
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
entrypoint: ["/app/llama-server"]
command:
- --model
- /models/${CHARRP_MODEL:-rp/TheDrummer_Magidonia-24B-v4.3-Q6_K.gguf}
- --host
- 0.0.0.0
- --port
- "8080"
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_CTX:-98304}"
- --flash-attn
- on
# q8_0 KV cache ~halves KV VRAM (8-bit, near-lossless) → ~2x the context per GB.
# Mistral/Magistral handles q8 KV cleanly. Set f16 in .env to disable.
- --cache-type-k
- ${CHARRP_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_KV_TYPE:-q8_0}
- --jinja
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 240s
networks:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=char-rp (Magidonia-24B GGUF)
- homepage.icon=mdi-drama-masks
- homepage.description=Dark-romantasy RP prose seat, non-thinking (llama.cpp, fv-ml1 GPU 0)
- homepage.href=http://10.251.50.54:${CHARRP_PORT:-8016}
# ── REASONING seat — NEO-CODE = Heretic2-Thinking (Qwen3.6-27B) managed thinking. gateway char-rp-reasoning. ──
llama-charrp-reasoning:
# ⚠️ CUSTOM llama.cpp build (master 6eddde0 + unmerged PR #25544). Needed for two reasons:
# (1) recent master parses Qwen3.6's native qwen3_coder tool-call format (stock b8840 predates it —
# the <tool_call><function=..><parameter=..> XML is Qwen3.5/3.6-native, NOT an OpenHands quirk);
# (2) PR #25544 multi-terminator reasoning-budget fix (Worldtree #355) — belt-and-suspenders now that
# NEO-CODE shows 0.0 runaway (R36 gate), but keep it. DO NOT revert to stock until #25544 merges.
# Build recipe + why + rollback: ./llamacpp-custom/README.md.
# Rollback: set LLAMA_REASONING_IMAGE=ghcr.io/mostlygeek/llama-swap:cuda in .env + recreate.
image: ${LLAMA_REASONING_IMAGE:-llamacpp-charrp:custom-latest}
container_name: ${CHARRP_REASONING_CONTAINER:-llama-charrp-reasoning}
restart: unless-stopped
runtime: nvidia
# Sequence AFTER the prose seat is healthy so the two GPU-0 allocations don't race.
depends_on:
llama-charrp:
condition: service_healthy
ports:
- "${CHARRP_REASONING_PORT:-8018}:8080"
volumes:
- ${MODELS_DIR:-/tank/aimodels/llm}:/models:ro
environment:
- NVIDIA_VISIBLE_DEVICES=${CHARRP_GPU_ID:-0}
entrypoint: ["/app/llama-server"]
command:
- --model
- /models/${CHARRP_REASONING_MODEL:-rp/Qwen3.6-27B-NEO-CODE-HERE-2T-OT-Q5_K_M.gguf}
- --host
- 0.0.0.0
- --port
- "8080"
- --n-gpu-layers
- "999"
- --ctx-size
- "${CHARRP_REASONING_CTX:-40960}"
- --flash-attn
- on
# NEO-CODE = Qwen3.6-27B GDN-hybrid (16 of 64 layers cache KV → KV cheap); native ctx 262144
# (256K). Full 256K @ q8_0 KV ≈ 8.6G, fits GPU0 w/ ~3.8G margin. q8_0 coherent; f16 in .env if gibberish.
- --cache-type-k
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --cache-type-v
- ${CHARRP_REASONING_KV_TYPE:-q8_0}
- --jinja
# NEO-CODE's Qwen3.6 template natively opens <think> → llama.cpp manages the reasoning
# (trace to reasoning_content, content stays clean prose); --reasoning-budget caps the CoT.
# (R36 gate 2026-07-14: NEO-CODE composite 0.922 tool-calling + 0.0 runaway — beat Deckard
# 0.08/0.80 and gen-reasoning 0.856. Budget held at 400: latency-coupled to soong's client timeout.)
- --reasoning
- on
- --reasoning-format
- deepseek
- --reasoning-budget
- "${CHARRP_REASONING_BUDGET:-400}"
# Sampler defaults per the DavidAU/Qwen3.6 model card (thinking-mode, general tasks): temp 1.0,
# top_p 0.95, top_k 20, min_p 0.0, no rep-penalty, no DRY (DRY was a QwQ/Deckard looping band-aid
# NEO-CODE doesn't need). All tunable via .env. NOTE: 0.922 tool-gate was on the OLD Deckard
# samplers (effective temp~0.8 + DRY); re-validate tools + slop on these card samplers.
- --temp
- "${CHARRP_REASONING_TEMP:-1.0}"
- --top-p
- "${CHARRP_REASONING_TOP_P:-0.95}"
- --top-k
- "${CHARRP_REASONING_TOP_K:-20}"
- --min-p
- "${CHARRP_REASONING_MIN_P:-0.0}"
healthcheck:
test: ["CMD-SHELL", "curl -fsS http://localhost:8080/health >/dev/null || exit 1"]
interval: 30s
timeout: 10s
retries: 3
start_period: 300s
networks:
- tnet
labels:
- homepage.group=AI - Dormant
- homepage.name=char-rp-reasoning (QwQ-32B RpR-v4 GGUF)
- homepage.icon=mdi-brain
- homepage.description=Dark-romantasy RP reasoning seat, managed CoT (llama.cpp, fv-ml1 GPU 0)
- homepage.href=http://10.251.50.54:${CHARRP_REASONING_PORT:-8018}
networks:
tnet:
name: traefik-net
external: true