feat(litellm): wire canonical sampler defaults for all 4 gateway seats
dvalin-smithy canonical set, infra-ops triaged + char-rp A/B-validated on the live serve. - gen (+summarizer-large twin): presence_penalty 1.0 -> 1.5 (Qwen3.6 non-thinking rec). - gen-reasoning: temp 0.6 -> 1.0, presence 1.0 -> 1.5 (Qwen general-thinking profile; the old 0.6 was the coding sub-profile). - char-rp: temp 1.0 -> 1.1, min_p 0.03 -> 0.10, top_k 0, NO rep. A/B on 2 dark-romantasy prompts: min_p 0.10 richened imagery; repeat_penalty 1.05 REJECTED (injected a stray markdown title, hurts Drummer/Magistral RP creativity per the card + dvalin's own note). - char-rp-reasoning: add explicit top_p 0.95 (else per the RpR card: no rep/DRY/XTC). Canonical reference: docs/pfi/model-sampler-defaults.md (mirrors dvalin's derivation).
This commit is contained in:
@@ -62,7 +62,7 @@ model_list:
|
||||
model: hosted_vllm/qwen3.6-27b-aeon
|
||||
api_base: http://10.250.50.54:8015/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
presence_penalty: 1.0
|
||||
presence_penalty: 1.5
|
||||
temperature: 0.7
|
||||
top_p: 0.8
|
||||
extra_body:
|
||||
@@ -115,15 +115,16 @@ model_list:
|
||||
# RETIRED with the displacement (→ 404, callers migrate to gen): qwen3.5-122-a10b
|
||||
# [-reasoning] + qwen-large[-reasoning] — they named a 122B that no longer exists;
|
||||
# aliasing a 27B under those is the naming footgun the qwen36-vl stack warns against.
|
||||
# presence_penalty: 1.0 INHERITED from qwopus (same-family abliterated/NVFP4 anti-
|
||||
# repetition damper, operator 2026-06-27) — RE-VALIDATE for AEON; NOT yet confirmed
|
||||
# for this model's repetition behavior. ---
|
||||
# presence_penalty: 1.5 — Qwen3.6 README anti-repetition rec for BOTH non-thinking and
|
||||
# thinking (dvalin-smithy canonical 2026-07-08, validated vs Qwen guidance). gen
|
||||
# non-thinking temp 0.7/top_p 0.8; gen-reasoning thinking temp 1.0/top_p 0.95 (the
|
||||
# GENERAL thinking profile, not the 0.6 coding sub-profile). docs/pfi/model-sampler-defaults.md. ---
|
||||
- model_name: gen
|
||||
litellm_params:
|
||||
model: hosted_vllm/qwen3.6-27b-aeon
|
||||
api_base: http://10.250.50.54:8015/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
presence_penalty: 1.0
|
||||
presence_penalty: 1.5
|
||||
temperature: 0.7
|
||||
top_p: 0.8
|
||||
extra_body:
|
||||
@@ -139,8 +140,8 @@ model_list:
|
||||
model: hosted_vllm/qwen3.6-27b-aeon-thinking
|
||||
api_base: http://10.250.50.54:8015/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
presence_penalty: 1.0
|
||||
temperature: 0.6
|
||||
presence_penalty: 1.5
|
||||
temperature: 1.0
|
||||
top_p: 0.95
|
||||
extra_body:
|
||||
top_k: 20
|
||||
@@ -152,19 +153,23 @@ model_list:
|
||||
# ana-ml2 GPU 0). TheDrummer Magidonia-24B-v4.3 Q6_K — Magistral (Mistral) RP tune.
|
||||
# NON-thinking: elite literary prose, zero refusal on dark/explicit scenes, ~65 tok/s,
|
||||
# tight POV/instruction adherence (live-tested 2026-07-08). Replaced the broken Angel
|
||||
# NVFP4 serve AND the earlier AEON-rp MTP twin. Magistral is stable WITHOUT a repetition
|
||||
# penalty (dropped the old 1.05); min_p 0.03 is the anti-slop knob. No enable_thinking
|
||||
# kwarg — meaningless to the Mistral template. Alt prose model (swap via the stack .env):
|
||||
# MS3.2-PaintedFantasy-v4.1-24B. Callers may override the sampler.
|
||||
# NVFP4 serve AND the earlier AEON-rp MTP twin. Sampler A/B-tuned 2026-07-08 vs the
|
||||
# dvalin-smithy canonical: temp 1.1 / top_p 0.95 / min_p 0.10 / top_k 0, NO repetition
|
||||
# penalty. min_p 0.10 richened imagery vs 0.03; rep 1.05 REJECTED (injected a stray
|
||||
# markdown title in a grief scene — rep-style penalties hurt Drummer RP, matching the
|
||||
# card). No enable_thinking kwarg — meaningless to the Mistral template. Alt prose model
|
||||
# (swap via the stack .env): MS3.2-PaintedFantasy-v4.1-24B. Callers may override.
|
||||
# docs/pfi/model-sampler-defaults.md.
|
||||
- model_name: char-rp
|
||||
litellm_params:
|
||||
model: hosted_vllm/magidonia-24b-v4.3
|
||||
api_base: http://10.250.50.54:8016/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
temperature: 1.0
|
||||
temperature: 1.1
|
||||
top_p: 0.95
|
||||
extra_body:
|
||||
min_p: 0.03
|
||||
min_p: 0.10
|
||||
top_k: 0
|
||||
model_info:
|
||||
mode: chat
|
||||
# char-rp-reasoning -> GGUF managed-REASONING seat (:8018, llama.cpp, char-rp-gguf stack).
|
||||
@@ -181,6 +186,7 @@ model_list:
|
||||
api_base: http://10.250.50.54:8018/v1
|
||||
api_key: os.environ/VLLM_API_KEY
|
||||
temperature: 1.0
|
||||
top_p: 0.95
|
||||
extra_body:
|
||||
top_k: 40
|
||||
min_p: 0.02
|
||||
|
||||
Reference in New Issue
Block a user