feat(reward-seat): move Skywork reward seat from fv-ml1 to esh-ml1; audit finds nothing superseding it

- Audit: Skywork-Reward-V2-Llama-3.1-8B is still #1 of 188 on AllenAI's
  RewardBench 2 per-sample results; no Skywork V3; the -40M sibling is
  vendor-marked experimental. Our AWQ W4A16 quant: 0.847 vs published bf16
  0.860 on a 150-prompt sample (within +/-2.9 pt SE), 96.2% pairwise
  agreement. Double BOS from vLLM on pre-templated text costs a further
  ~2.7 pts; callers must send add_special_tokens=false.
- No working consumer: 0 requests since 2026-09-13; Worldtree Domari points
  at a dead IP with a non-vLLM schema (reported to worldtree-dev).
- Move: sha256-identical model copy; vLLM v0.24.0 on esh-ml1 :8003 at 0.55
  util (KV 1.30x of a 16k request). Parity vs fv-ml1: 149/150 verdicts,
  99.8% pairwise signs, raw |delta| median 0.049.
- Gateway /scalar-judge passthrough -> 10.0.50.80:8003; fv-ml1 vllm-reward
  removed (~10.2 GB freed on GPU 1). stacks/vllm now holds only vllm-coder.
This commit is contained in:
vh
2026-09-25 09:09:48 -07:00
parent 7bdac80878
commit 52612cbe96
8 changed files with 207 additions and 72 deletions
+10 -62
View File
@@ -1,4 +1,4 @@
# vLLM — utility seats on fv-ml1: Skywork Reward-V2 classifier + the coder FIM seat.
# vLLM — the coder FIM seat on fv-ml1 (the last utility seat left here).
#
# Originally created to replace the unmaintained Infinity stack (embed +
# rerank); generalized 2026-05-13 to host any vLLM-served model on fv-ml1,
@@ -7,13 +7,14 @@
#
# ⚠ Embedding + reranking LEFT this stack 2026-09-25 (Prime): they now run on TEI
# on esh-ml1 (stacks/embed-rerank), and TEI is the fleet's embed/rerank engine
# from now on. Do not re-add them here.
# from now on. The reward classifier moved to esh-ml1 the same day
# (stacks/reward-seat). Do not re-add them here.
#
# All tunables live in .env — edit that, not this file.
#
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on fv-ml1 and is
# bind-mounted into the reward service at /local-models. Not from HF Hub.
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model. The canonical
# copy still lives at /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on
# fv-ml1; the serving copy is on esh-ml1. Not from HF Hub.
services:
# vllm-embed (Qwen3-Embedding-0.6B, :8001) and vllm-rerank-a3
@@ -23,63 +24,10 @@ services:
# and `reranker-a3-bge-v2-m3` did not change. Their definitions are in git
# history before this commit if a rollback is ever needed.
vllm-reward:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-reward
restart: unless-stopped
ipc: host
ports:
- "${REWARD_PORT}:8000"
volumes:
# AWQ output lives in the legacy llama-swap models tree, not the HF cache
# — bind-mount the LLM models dir read-only so the reward service can
# load it as a local-path HF-format model.
- /tank/aimodels/llm:/local-models:ro
environment:
- VLLM_API_KEY=${API_KEY:-}
command:
- /local-models/Skywork-Reward-V2-Llama-3.1-8B-AWQ
- --served-model-name
- Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
# vLLM 0.19.1 deprecated --task in favor of --runner. The model's
# config.json declares `LlamaForSequenceClassification` so the
# pooling runner uses it as a classifier (single-label reward score)
# without needing an explicit task flag.
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${REWARD_GPU_MEM_UTIL}
- --max-model-len
- ${REWARD_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 240s
networks:
- tnet
labels:
- homepage.group=AI - Eval & Retrieval
- homepage.name=vLLM Reward (Skywork)
- homepage.icon=mdi-scale-balance
- homepage.description=Skywork-Reward-V2 8B classifier via vLLM (fv-ml1)
- homepage.href=http://10.251.50.54:${REWARD_PORT}/docs
# vllm-reward (Skywork-Reward-V2-Llama-3.1-8B AWQ, :8003) — MOVED to esh-ml1
# 2026-09-25 (stacks/reward-seat); the gateway `/scalar-judge` passthrough
# follows it. The model files stay at /tank/aimodels/llm/ on fv-ml1 as the
# canonical copy of the local quant.
# vllm-granite (ibm-granite/granite-4.1-8b-fp8, :8004) — RETIRED 2026-08-12,
# service block removed 2026-08-20. It was the fleet summarizer until the