feat(reward-seat): move Skywork reward seat from fv-ml1 to esh-ml1; audit finds nothing superseding it
- Audit: Skywork-Reward-V2-Llama-3.1-8B is still #1 of 188 on AllenAI's RewardBench 2 per-sample results; no Skywork V3; the -40M sibling is vendor-marked experimental. Our AWQ W4A16 quant: 0.847 vs published bf16 0.860 on a 150-prompt sample (within +/-2.9 pt SE), 96.2% pairwise agreement. Double BOS from vLLM on pre-templated text costs a further ~2.7 pts; callers must send add_special_tokens=false. - No working consumer: 0 requests since 2026-09-13; Worldtree Domari points at a dead IP with a non-vLLM schema (reported to worldtree-dev). - Move: sha256-identical model copy; vLLM v0.24.0 on esh-ml1 :8003 at 0.55 util (KV 1.30x of a 16k request). Parity vs fv-ml1: 149/150 verdicts, 99.8% pairwise signs, raw |delta| median 0.049. - Gateway /scalar-judge passthrough -> 10.0.50.80:8003; fv-ml1 vllm-reward removed (~10.2 GB freed on GPU 1). stacks/vllm now holds only vllm-coder.
This commit is contained in:
@@ -1047,15 +1047,19 @@ general_settings:
|
||||
# unbounded. Names verified against LiteLLM docs (proxy/spend_logs_deletion).
|
||||
maximum_spend_logs_retention_period: "7d"
|
||||
maximum_spend_logs_retention_interval: "1d"
|
||||
# scalar-judge → Skywork-Reward-V2 (scalar reward model; vLLM pooling on
|
||||
# fv-ml1:8003). LiteLLM has no reward/pooling MODE, so this is a passthrough,
|
||||
# scalar-judge → Skywork-Reward-V2-Llama-3.1-8B (AWQ; vLLM pooling on esh-ml1
|
||||
# :8003, moved from fv-ml1 2026-09-25, same files, verdicts identical 149/150).
|
||||
# LiteLLM has no reward/pooling MODE, so this is a passthrough,
|
||||
# not a model_list alias. Gateway-key-gated. Consumers POST the reward body to
|
||||
# /scalar-judge/<route> (e.g. /pooling or /classify), forwarded to :8003.
|
||||
# /scalar-judge/<route> (e.g. /classify), forwarded to :8003.
|
||||
# SWAP-SENSITIVE: a different reward model shifts the score scale, so consumers
|
||||
# must recalibrate thresholds after a backing swap.
|
||||
# ⚠ Send already-templated text with "add_special_tokens": false. Otherwise vLLM
|
||||
# prepends a SECOND <|begin_of_text|>, which cost ~2.7 pts on RewardBench 2
|
||||
# (stacks/reward-seat/README.md).
|
||||
pass_through_endpoints:
|
||||
- path: "/scalar-judge"
|
||||
target: "http://10.251.50.54:8003"
|
||||
target: "http://10.0.50.80:8003"
|
||||
forward_headers: true
|
||||
include_subpath: true
|
||||
|
||||
|
||||
Reference in New Issue
Block a user