feat(litellm): wire char-rp-reasoning + char-rp-fable to the Fable-Fusion seat

Operator-directed evaluation window. char-rp-reasoning now resolves to
Fable-Fusion 711 on :8019 instead of Dark-Scarlett on :8018; DS is DOWN
because GPU1 is zero-sum and FF occupies her slot.

This is an EXPLICIT substitution, not a silent alias swap: the config
block says so in place, carries the measured justification, and names the
rollback. char-rp-fable is added as the seat's honest name so the
evaluation can address it without depending on the temporary repoint, and
as a distinct model_name it gets its own litellm_params object rather than
sharing one (which is what bleeds sampler overrides between variants).

Samplers are unchanged from the DS entry and match the model card's
thinking-mode recommendation (temp 1.0 / top_p 0.95 / top_k 20). Verified
the FF chat template actually honours enable_thinking
(chat_template.jinja:44) rather than ignoring it -- the mismatch that
returned null content on the MeroMero seat.

Verified end-to-end through the gateway on both aliases: prose in content,
CoT in reasoning_content, finish=stop.

CONSUMER HAZARD: FF reasons heavily (2.1-4.6k chars). At max_tokens=1200
one of seven calls returned EMPTY content with finish_reason=length --
reasoning ate the whole budget. Not a refusal and not an alias fault. Use
max_tokens >= 3072; 6/6 clean there. No default is baked into the alias
because that would override caller intent silently.
This commit is contained in:
2026-08-16 09:28:24 -07:00
parent b9e68c3fd2
commit ee2b678bcb
+44 -2
View File
@@ -194,10 +194,52 @@ model_list:
# NO presence/rep penalty; DRY 0.8 server-side). Tuning ladder: flat prose→min_p 0.08, loops→DRY 0.9,
# over-damped→DRY 0.6/off. Do NOT import RpR/QwQ sampler rules (different family). NOT the same model as
# char-rp (best-of-breed per seat) — see stacks/char-rp-gguf/README.md.
# ⚠️ TEMPORARY REPOINT 2026-08-16 (operator-directed evaluation window).
# char-rp-reasoning currently resolves to FABLE-FUSION 711 on :8019, NOT to
# Dark-Scarlett. DS v1.0 is DOWN — GPU1 is zero-sum and Fable-Fusion occupies
# her slot. This is a deliberate, explicit substitution for hands-on testing;
# it is NOT a silent alias swap, and it is not the permanent seat decision.
# Address the seat unambiguously as `char-rp-fable` below; `char-rp-reasoning`
# is kept live only so existing consumers keep working during the window.
#
# WHY: DS v1.0 is a plain finetune of stock Qwen3.6-27B with NO abliteration,
# so cold prompts revert to safety-tuned base behaviour. Measured three-arm
# A/B (services/refusal-probe/): under a bare instruction with no character
# card, DS refuses 92.5% (37/40) and Fable-Fusion 15.8% (6/38); with a
# character card both sit at ~0%. Fable-Fusion is Heretic-abliterated.
#
# ROLLBACK (restores Dark-Scarlett):
# ssh infra-ops@10.250.50.54 'cd /opt/docker/compose/fablefusion-charrp-probe && sudo docker compose down'
# ssh infra-ops@10.250.50.54 'cd /opt/docker/compose/darkscarlett-charrp-reasoning && sudo docker compose up -d'
# then revert this block to api_base :8018 / model hosted_vllm/char-rp-reasoning
# and restart litellm (~52s).
#
# Samplers below are the model card's thinking-mode recommendation (temp 1.0 /
# top_p 0.95 / top_k 20) and are unchanged from the DS entry. Verified the FF
# chat template honours `enable_thinking` (chat_template.jinja:44) rather than
# ignoring it — the mismatch that returned null content on the MeroMero seat.
- model_name: char-rp-reasoning
litellm_params:
model: hosted_vllm/char-rp-reasoning
api_base: http://10.250.50.54:8018/v1
model: hosted_vllm/char-rp-probe
api_base: http://10.250.50.54:8019/v1
api_key: os.environ/VLLM_API_KEY
temperature: 1.0
top_p: 0.95
extra_body:
top_k: 20
chat_template_kwargs:
enable_thinking: true
model_info:
mode: chat
# char-rp-fable -> the SAME Fable-Fusion seat under its own honest name, so the
# evaluation can address it without relying on the temporary repoint above.
# Distinct model_name = distinct litellm_params object, which avoids the
# shared-deployment param mutation that bleeds sampler overrides between
# variants.
- model_name: char-rp-fable
litellm_params:
model: hosted_vllm/char-rp-probe
api_base: http://10.250.50.54:8019/v1
api_key: os.environ/VLLM_API_KEY
temperature: 1.0
top_p: 0.95