vllm: rename stack from vllm-qwen3 → vllm + add Skywork reward classifier

Two related changes shipped together. The stack rename is independent
but adding `vllm-reward` to the existing `vllm-qwen3` would have made
that name actively misleading.

**Rename:** `stacks/vllm-qwen3/ → stacks/vllm/`. Updated all in-repo
references (README.md root, servers/ana-ml2/, stacks/llama-swap/,
configs/restic/ana-ml2/, docs/runbooks/disaster-recovery.md). Two
intentional history mentions retained (servers/ana-ml2 + stacks/vllm
README).

**Add `vllm-reward` service:** serves Skywork-Reward-V2-Llama-3.1-8B-AWQ
on port 8003. The AWQ output is a locally-quantized model (not from HF),
so bind-mounts `/tank/aimodels/llm:/local-models:ro` rather than the
shared HF cache. Model config.json declares LlamaForSequenceClassification
which vLLM's pooling runner picks up automatically — produces a single
reward score per input via /classify.

**Flag note:** the user's spec listed `--task classify`, but vLLM 0.19.1
deprecated --task in favor of --runner pooling (model architecture in
config.json drives the classification head). Compose uses --runner
pooling with a comment explaining the substitution.

**GPU memory:** no rebalance needed — production had already tuned
EMBED/RERANK down from 0.40 to 0.20 each (canonical .env.example now
matches reality). Adding REWARD at 0.30 totals 0.70, leaving ~14 GB
headroom on the 48 GB Ada.

**Server-side:** brought existing vllm-qwen3 down, mv'd
/opt/docker/compose/vllm-qwen3 → /opt/docker/compose/vllm, appended
REWARD_* lines to existing .env (preserving API_KEY/HF_TOKEN), deployed
new compose via scripts/deploy-stack.sh, brought all 3 services up.

**Smoke tests:**
- /health on 8001/8002/8003 → 200
- /v1/models on 8003 → lists Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
  with max_model_len 16384
- /classify with a sample conversation → returns LABEL_0 with prob 0.9999
  (single-output regression-style reward score, expected shape for a
  reward model)
This commit is contained in:
vh
2026-05-13 22:00:26 -07:00
parent 662a73ee0e
commit 7e7130172e
12 changed files with 428 additions and 135 deletions
+1 -1
View File
@@ -12,7 +12,7 @@ GGUF model server with on-demand model swapping. Served via llama.cpp's `llama-s
- **`.env.example`** — template for the per-host `.env`. Copy to `.env` on the server and tweak.
- **`config.yaml`** — model definitions and groups. Deployed to `/opt/docker/conf/llama-swap/config.yaml` on the server.
Homepage labels are in the compose file under the `AI Systems` group, matching the convention used by `vllm-qwen3` and `infinity`.
Homepage labels are in the compose file under the `AI Systems` group, matching the convention used by `vllm` and `infinity`.
## Deploy a fresh install
-40
View File
@@ -1,40 +0,0 @@
# vllm-qwen3 stack tunables. Copy this to `.env` on the server before deploying.
#
# cp .env.example .env
# # edit .env with real values
# docker compose up -d
# Image version — pin for reproducibility (`latest` for edge)
VLLM_VERSION=latest
# Host ports (container always listens on 8000 internally)
EMBED_PORT=8001
RERANK_PORT=8002
# GPU assignment — both services share this GPU
# (ana-ml2 has 0 and 1; default 1 keeps 0 free for heavy LLM work)
GPU_ID=1
# Models — reference by full repo name in API requests
EMBED_MODEL=Qwen/Qwen3-Embedding-0.6B
RERANK_MODEL=Qwen/Qwen3-Reranker-0.6B
# GPU memory split — fractions are of TOTAL GPU memory, not free memory.
# When two vLLM services share a GPU, each profiler needs its own slice to
# fit both the model and KV cache, so small values cause the second-to-start
# service to OOM on KV cache allocation. 0.40 + 0.40 leaves ~20% headroom
# and is comfortably above the minimum for two 0.6B Qwen3 models at 8k ctx.
EMBED_GPU_MEM_UTIL=0.40
RERANK_GPU_MEM_UTIL=0.40
# Context length caps — lower these if VRAM is tight.
# Qwen3-Embedding supports up to 32k; reranker up to 32k.
EMBED_MAX_MODEL_LEN=8192
RERANK_MAX_MODEL_LEN=8192
# Optional API key — leave blank for no auth (fine on the internal network).
# If set, both services require `Authorization: Bearer <key>`.
API_KEY=
# HuggingFace token — only needed for gated models
HF_TOKEN=
-80
View File
@@ -1,80 +0,0 @@
# vllm-qwen3
Qwen3 embedding + reranker served via vLLM. Replaces the unmaintained Infinity stack.
**Server:** ana-ml2
**Ports:** `8001` (embed), `8002` (rerank) — both configurable via `.env`
**GPU:** both services share GPU 1 by default (configurable)
## Why two services
vLLM runs **one model per process**, so embedding and reranking each get their own container. Both pin to the same GPU and split VRAM via `--gpu-memory-utilization`. Both use `--runner pooling` so the OpenAI server exposes `/v1/embeddings` (for the embedder) and `/rerank`, `/score` (for the reranker, which also needs the `--hf-overrides` described below).
## Reranker caveat
Qwen/Qwen3-Reranker-0.6B is a causal-LM checkpoint. The `--hf-overrides` flag in `compose.yaml` re-maps it to `Qwen3ForSequenceClassification` so vLLM's `/rerank` and `/score` endpoints work and the model emits only `no`/`yes` class logits instead of the full 151k-token distribution.
If that override breaks after a vLLM upgrade, the pre-converted checkpoint `tomaarsen/Qwen3-Reranker-0.6B-seq-cls` is a drop-in replacement that needs no overrides — set `RERANK_MODEL=tomaarsen/Qwen3-Reranker-0.6B-seq-cls` in `.env` and remove the `--hf-overrides` line from the compose.
## Deploy
```bash
# On ana-ml2:
sudo mkdir -p /opt/docker/compose/vllm-qwen3
sudo chown $USER /opt/docker/compose/vllm-qwen3
cd /opt/docker/compose/vllm-qwen3
# Copy compose.yaml + .env.example here (e.g. via scp from this workspace)
cp .env.example .env
# edit .env — pick GPU, ports, memory split, etc.
# Pre-download models (optional, speeds first boot)
HF_HOME=/tank/aimodels/huggingface hf download "$(grep ^EMBED_MODEL .env | cut -d= -f2)"
HF_HOME=/tank/aimodels/huggingface hf download "$(grep ^RERANK_MODEL .env | cut -d= -f2)"
# Dry-parse
docker compose config
# Launch
docker compose up -d
docker compose logs -f
```
First boot compiles CUDA graphs and can take 2–3 minutes per service. The `start_period: 180s` healthcheck grace reflects that.
## Verify
```bash
# Health
curl -s http://localhost:8001/health
curl -s http://localhost:8002/health
# Embedding (OpenAI-compatible)
curl -s http://localhost:8001/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-Embedding-0.6B","input":["hello world"]}' | jq .
# Reranker
curl -s http://localhost:8002/rerank \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-Reranker-0.6B","query":"what is a cat","documents":["cats are mammals","dogs bark"]}' | jq .
# Listed models
curl -s http://localhost:8001/v1/models | jq .
curl -s http://localhost:8002/v1/models | jq .
```
## Scaling knobs
- **`EMBED_GPU_MEM_UTIL` / `RERANK_GPU_MEM_UTIL`** — fractions of **total** GPU VRAM each service reserves (not of free VRAM). Both services profile independently, so each slice must be large enough to fit that service's model + KV cache with no knowledge of the other. Setting them too low causes the second-to-start container to OOM on KV cache allocation with `Available KV cache memory: -X.XX GiB`. Default 0.40/0.40 (= 0.80 total) leaves ~20% GPU headroom and works cleanly for the 0.6B pair at 8k context; raise for 4B/8B variants or drop `max-model-len` if you need more room.
- **`EMBED_MAX_MODEL_LEN` / `RERANK_MAX_MODEL_LEN`** — lower to reduce KV-cache allocation if VRAM is tight. Qwen3 supports up to 32k natively.
- **Larger models** — Qwen3-Embedding/Reranker come in 0.6B / 4B / 8B. Swap `EMBED_MODEL` / `RERANK_MODEL` and bump the memory fractions accordingly.
- **Separate GPUs** — if contention hurts latency, split them: add a second `GPU_ID_RERANK` variable and point each service at its own device. (Requires a small compose edit; currently both share `${GPU_ID}`.)
## Migrating off Infinity
Once this stack is verified stable:
1. Stop the infinity stack (`docker compose down` under `/opt/docker/compose/infinity/`).
2. Update consumers (AIPA agents, LibreChat RAG) to point at `:8001` for embeddings and `:8002` for rerank.
3. Delete `stacks/infinity/` from this workspace.
+54
View File
@@ -0,0 +1,54 @@
# vllm stack tunables. Copy this to `.env` on the server before deploying.
#
# cp .env.example .env
# # edit .env with real values
# docker compose up -d
# Image version — pin for reproducibility (`latest` for edge)
VLLM_VERSION=latest
# Host ports (container always listens on 8000 internally)
EMBED_PORT=8001
RERANK_PORT=8002
REWARD_PORT=8003
# GPU assignment — all services share this GPU
# (ana-ml2 has 0 and 1; default 1 keeps 0 free for heavy LLM work in llama-swap)
GPU_ID=1
# Models — reference by full repo name in API requests
EMBED_MODEL=Qwen/Qwen3-Embedding-0.6B
RERANK_MODEL=Qwen/Qwen3-Reranker-0.6B
# Skywork is a local-path AWQ output, not from HF Hub. Bind-mounted into the
# reward container at /local-models — see compose.yaml. No env var here for
# the model path itself since it's hard-coded in the compose command.
# GPU memory split — fractions are of TOTAL GPU memory, not free memory.
# vLLM profiles each service independently, so each slice must be large enough
# to fit that service's model + KV cache with no awareness of the others.
# Setting any one too low causes that container to OOM on KV cache allocation
# with `Available KV cache memory: -X.XX GiB`.
#
# Layout on a 48 GB Ada (~30% headroom kept free; matches current production):
# EMBED 0.20 (~9.6 GB) — 0.6B Qwen3 embed at 8k ctx; comfortable
# RERANK 0.20 (~9.6 GB) — 0.6B Qwen3 rerank at 8k ctx; comfortable
# REWARD 0.30 (~14 GB) — 8B Skywork AWQ at 16k ctx classify; tune up if OOM
EMBED_GPU_MEM_UTIL=0.20
RERANK_GPU_MEM_UTIL=0.20
REWARD_GPU_MEM_UTIL=0.30
# Context length caps — lower these if VRAM is tight.
# Qwen3-Embedding supports up to 32k; reranker up to 32k.
# Skywork capped at 16k server-side as defense-in-depth; JudgeClient also
# enforces the cap at dispatch time per spec.
EMBED_MAX_MODEL_LEN=8192
RERANK_MAX_MODEL_LEN=8192
REWARD_MAX_MODEL_LEN=16384
# Optional API key — leave blank for no auth (fine on the internal network).
# If set, all three services require `Authorization: Bearer <key>`.
API_KEY=
# HuggingFace token — only needed for gated models in the HF-Hub-loaded
# services (embed/rerank). Reward is local-path, ignores this.
HF_TOKEN=
+133
View File
@@ -0,0 +1,133 @@
# vllm
Multi-service vLLM stack on ana-ml2. Started as Qwen3 embedding + rerank
(replacing the unmaintained Infinity stack); generalized to host any vLLM
model on the box, currently three services:
- **`vllm-embed`** — Qwen3-Embedding 0.6B → OpenAI `/v1/embeddings`
- **`vllm-rerank`** — Qwen3-Reranker 0.6B → `/rerank` + `/score`
- **`vllm-reward`** — Skywork-Reward-V2-Llama-3.1-8B-AWQ → `/classify`
**Server:** ana-ml2
**Ports:** `8001` (embed), `8002` (rerank), `8003` (reward) — all configurable via `.env`
**GPU:** all three services share GPU 1 by default (configurable)
## Why one process per service
vLLM runs **one model per process**, so each model gets its own container.
All three pin to the same GPU and split VRAM via `--gpu-memory-utilization`.
Embed + rerank use `--runner pooling`; reward uses `--task classify` (newer
vLLM flag for sequence-classification heads).
## Reranker caveat
Qwen/Qwen3-Reranker-0.6B is a causal-LM checkpoint. The `--hf-overrides`
flag in `compose.yaml` re-maps it to `Qwen3ForSequenceClassification` so
vLLM's `/rerank` and `/score` endpoints work and the model emits only
`no`/`yes` class logits instead of the full 151k-token distribution.
If that override breaks after a vLLM upgrade, the pre-converted checkpoint
`tomaarsen/Qwen3-Reranker-0.6B-seq-cls` is a drop-in replacement that needs
no overrides — set `RERANK_MODEL=tomaarsen/Qwen3-Reranker-0.6B-seq-cls` in
`.env` and remove the `--hf-overrides` line from the compose.
## Reward / Skywork specifics
`vllm-reward` serves `Skywork-Reward-V2-Llama-3.1-8B-AWQ` — an AWQ
quantization produced locally (not pulled from HF). The quant output lives
at `/tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ` on ana-ml2 and
is bind-mounted read-only into the container at `/local-models`. vLLM
loads it as a local-path HF-format model (config.json + safetensors).
`--max-model-len 16384` is a server-side cap; JudgeClient on the consumer
side also enforces this at dispatch time. Defense-in-depth.
`--dtype auto` lets vLLM pick the right path for AWQ-quantized weights.
## Deploy
```bash
# Sync canonical → ana-ml2
scripts/deploy-stack.sh ana-ml2 vllm
# Pre-download Qwen3 models from HF (Skywork is local-path, no pull needed)
scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
--var hf_repo=Qwen/Qwen3-Embedding-0.6B
scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
--var hf_repo=Qwen/Qwen3-Reranker-0.6B
# Launch
ssh ana-ml2 'cd /opt/docker/compose/vllm && docker compose up -d && docker compose logs --tail=30'
```
First boot compiles CUDA graphs and can take 2–3 minutes per service
(longer for the 8B reward — `start_period: 240s`).
## Verify
```bash
# Health (each on its own port)
curl -s http://localhost:8001/health
curl -s http://localhost:8002/health
curl -s http://localhost:8003/health
# Embedding (OpenAI-compatible)
curl -s http://localhost:8001/v1/embeddings \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-Embedding-0.6B","input":["hello world"]}' | jq .
# Reranker
curl -s http://localhost:8002/rerank \
-H "Content-Type: application/json" \
-d '{"model":"Qwen/Qwen3-Reranker-0.6B","query":"what is a cat","documents":["cats are mammals","dogs bark"]}' | jq .
# Reward classify (Skywork)
curl -s http://localhost:8003/classify \
-H "Content-Type: application/json" \
-d '{"model":"Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ","input":"User: hello\nAssistant: hi there"}' | jq .
# Listed models
curl -s http://localhost:8001/v1/models | jq .
curl -s http://localhost:8002/v1/models | jq .
curl -s http://localhost:8003/v1/models | jq .
```
## Scaling knobs
- **`EMBED_GPU_MEM_UTIL` / `RERANK_GPU_MEM_UTIL` / `REWARD_GPU_MEM_UTIL`** —
fractions of **total** GPU VRAM each service reserves (not of free VRAM).
Each profiler runs independently with no awareness of the others, so any
slice that's too small to fit `model + KV cache` will OOM the
second-to-start container with
`Available KV cache memory: -X.XX GiB`. Default 0.20/0.20/0.30 totals
0.70, leaving ~14 GB headroom on a 48 GB Ada (matches the production
tune-down from the original 0.40/0.40 defaults — embed/rerank fit
comfortably in 0.20 each). Bump REWARD up first if you see OOM — the
8B AWQ model's KV slice at 16k ctx is the tightest. Drop embed/rerank
further only if you've extended REWARD past 0.40 and still need room.
- **`EMBED_MAX_MODEL_LEN` / `RERANK_MAX_MODEL_LEN` / `REWARD_MAX_MODEL_LEN`** —
lower to reduce KV-cache allocation if VRAM is tight. Qwen3 supports up
to 32k natively; Skywork's reward cap of 16k is intentional (matches
JudgeClient's dispatch-side cap).
- **Larger Qwen3 models** — embedding/reranker come in 0.6B / 4B / 8B.
Swap `EMBED_MODEL` / `RERANK_MODEL` and bump the memory fractions
accordingly.
- **Separate GPUs** — if contention hurts latency, split them. Today all
three share `${GPU_ID}`. Adding `GPU_ID_REWARD`/etc. is a small compose
edit.
## History
- **2026-05-13 rename + extend** — stack renamed from `vllm-qwen3` to
`vllm` and gained the `vllm-reward` service (Skywork-Reward-V2 8B AWQ).
No GPU memory rebalance needed in practice — production had already
tuned EMBED/RERANK down from 0.40/0.40 to 0.20/0.20; adding REWARD at
0.30 fits cleanly with ~14 GB headroom on the 48 GB Ada.
- **Migrated off Infinity** — Infinity stopped shipping a `transformers`
build that knew Qwen3; this stack is the canonical replacement for both
embed and rerank. Consumers (AIPA agents, LibreChat RAG) point at
`:8001`/`:8002`.
@@ -1,10 +1,16 @@
# vLLM — Qwen3 Embedding + Reranker (one stack, two services).
# vLLM — Qwen3 Embedding + Reranker + Skywork Reward-V2 classifier.
#
# Replaces the unmaintained Infinity stack. vLLM runs one model per process,
# so this stack brings up two containers sharing a single GPU:
# Originally created to replace the unmaintained Infinity stack (embed +
# rerank); generalized 2026-05-13 to host any vLLM-served model on ana-ml2,
# starting with the Skywork-Reward-V2-Llama-3.1-8B reward classifier
# (AWQ-quantized locally, served from /tank/aimodels/llm/).
#
# vLLM runs one model per process, so this stack brings up three containers
# sharing a single GPU:
#
# vllm-embed — Qwen3-Embedding served as an OpenAI /v1/embeddings server
# vllm-rerank — Qwen3-Reranker served as a /rerank + /score server
# vllm-reward — Skywork-Reward-V2-Llama-3.1-8B-AWQ served as a /classify scorer
#
# The reranker is a causal-LM checkpoint; --hf-overrides re-maps it to
# Qwen3ForSequenceClassification so vLLM's reranking endpoints work and the
@@ -13,8 +19,14 @@
# All tunables live in .env — edit that, not this file.
#
# Pre-download models to avoid first-run delay:
# HF_HOME=/tank/aimodels/huggingface hf download Qwen/Qwen3-Embedding-0.6B
# HF_HOME=/tank/aimodels/huggingface hf download Qwen/Qwen3-Reranker-0.6B
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Embedding-0.6B
# scripts/elway ana-ml2 --playbook playbooks/pull-hf-repo.yaml \
# --var hf_repo=Qwen/Qwen3-Reranker-0.6B
#
# Skywork-Reward-V2-Llama-3.1-8B-AWQ is a locally-quantized model — lives at
# /tank/aimodels/llm/Skywork-Reward-V2-Llama-3.1-8B-AWQ on ana-ml2 and is
# bind-mounted into the reward service at /local-models. Not from HF Hub.
services:
vllm-embed:
@@ -127,6 +139,64 @@ services:
- homepage.description=Qwen3 Reranker via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${RERANK_PORT}/docs
vllm-reward:
image: vllm/vllm-openai:${VLLM_VERSION}
container_name: vllm-reward
restart: unless-stopped
ipc: host
ports:
- "${REWARD_PORT}:8000"
volumes:
# AWQ output lives in the legacy llama-swap models tree, not the HF cache
# — bind-mount the LLM models dir read-only so the reward service can
# load it as a local-path HF-format model.
- /tank/aimodels/llm:/local-models:ro
environment:
- VLLM_API_KEY=${API_KEY:-}
command:
- /local-models/Skywork-Reward-V2-Llama-3.1-8B-AWQ
- --served-model-name
- Skywork/Skywork-Reward-V2-Llama-3.1-8B-AWQ
# vLLM 0.19.1 deprecated --task in favor of --runner. The model's
# config.json declares `LlamaForSequenceClassification` so the
# pooling runner uses it as a classifier (single-label reward score)
# without needing an explicit task flag.
- --runner
- pooling
- --host
- 0.0.0.0
- --port
- "8000"
- --gpu-memory-utilization
- ${REWARD_GPU_MEM_UTIL}
- --max-model-len
- ${REWARD_MAX_MODEL_LEN}
- --dtype
- auto
deploy:
resources:
reservations:
devices:
- driver: nvidia
device_ids:
- "${GPU_ID}"
capabilities:
- gpu
healthcheck:
test: ["CMD", "curl", "-f", "http://localhost:8000/health"]
interval: 30s
timeout: 10s
retries: 3
start_period: 240s
networks:
- tnet
labels:
- homepage.group=AI Systems
- homepage.name=vLLM Reward (Skywork)
- homepage.icon=mdi-scale-balance
- homepage.description=Skywork-Reward-V2 8B classifier via vLLM (ana-ml2)
- homepage.href=http://10.250.50.54:${REWARD_PORT}/docs
networks:
tnet:
name: traefik-net