feat(reward-seat): move Skywork reward seat from fv-ml1 to esh-ml1; audit finds nothing superseding it

- Audit: Skywork-Reward-V2-Llama-3.1-8B is still #1 of 188 on AllenAI's
  RewardBench 2 per-sample results; no Skywork V3; the -40M sibling is
  vendor-marked experimental. Our AWQ W4A16 quant: 0.847 vs published bf16
  0.860 on a 150-prompt sample (within +/-2.9 pt SE), 96.2% pairwise
  agreement. Double BOS from vLLM on pre-templated text costs a further
  ~2.7 pts; callers must send add_special_tokens=false.
- No working consumer: 0 requests since 2026-09-13; Worldtree Domari points
  at a dead IP with a non-vLLM schema (reported to worldtree-dev).
- Move: sha256-identical model copy; vLLM v0.24.0 on esh-ml1 :8003 at 0.55
  util (KV 1.30x of a 16k request). Parity vs fv-ml1: 149/150 verdicts,
  99.8% pairwise signs, raw |delta| median 0.049.
- Gateway /scalar-judge passthrough -> 10.0.50.80:8003; fv-ml1 vllm-reward
  removed (~10.2 GB freed on GPU 1). stacks/vllm now holds only vllm-coder.
This commit is contained in:
vh
2026-09-25 09:09:48 -07:00
parent 7bdac80878
commit 52612cbe96
8 changed files with 207 additions and 72 deletions
+8 -4
View File
@@ -1047,15 +1047,19 @@ general_settings:
# unbounded. Names verified against LiteLLM docs (proxy/spend_logs_deletion).
maximum_spend_logs_retention_period: "7d"
maximum_spend_logs_retention_interval: "1d"
# scalar-judge → Skywork-Reward-V2 (scalar reward model; vLLM pooling on
# fv-ml1:8003). LiteLLM has no reward/pooling MODE, so this is a passthrough,
# scalar-judge → Skywork-Reward-V2-Llama-3.1-8B (AWQ; vLLM pooling on esh-ml1
# :8003, moved from fv-ml1 2026-09-25, same files, verdicts identical 149/150).
# LiteLLM has no reward/pooling MODE, so this is a passthrough,
# not a model_list alias. Gateway-key-gated. Consumers POST the reward body to
# /scalar-judge/<route> (e.g. /pooling or /classify), forwarded to :8003.
# /scalar-judge/<route> (e.g. /classify), forwarded to :8003.
# SWAP-SENSITIVE: a different reward model shifts the score scale, so consumers
# must recalibrate thresholds after a backing swap.
# ⚠ Send already-templated text with "add_special_tokens": false. Otherwise vLLM
# prepends a SECOND <|begin_of_text|>, which cost ~2.7 pts on RewardBench 2
# (stacks/reward-seat/README.md).
pass_through_endpoints:
- path: "/scalar-judge"
target: "http://10.251.50.54:8003"
target: "http://10.0.50.80:8003"
forward_headers: true
include_subpath: true