- services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py: linearize_moe first (playbook §3.15), asserts the expert Linear count, routers/vision/audio/norms/lm_head ignored, W4A16 for RP long-session fidelity, post-steps restore processor configs + template and reset the tokenizer truncation cap (§3.14); --dry-run proves targets before GPU time - services/erp-seat-quant/run_quant_erp_v6.sh: detached container on GPU1 (vllm-llmcompressor) - stacks/erp-seat: serve recipe copied from gemma4-charrp, true served name only, port 8021
13 lines
499 B
Bash
13 lines
499 B
Bash
# erp-seat — ana-ml2 GPU1. Real .env lives on the host at /opt/docker/compose/erp-seat/.env.
|
|
ERP_IMAGE=vllm/vllm-openai:v0.26.0
|
|
ERP_MODEL=/tank/aimodels/erp-tune-v6-nvfp4a16
|
|
ERP_SERVED_NAME=erp-tune-v6-nvfp4a16
|
|
ERP_CHAT_TEMPLATE=/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja
|
|
ERP_PORT=8021
|
|
ERP_GPU_ID=1
|
|
# 0.35 x 97.9 GiB = 34 GiB. GPU1 had ~47 GiB free on 2026-09-08 (scriberr/embed/rerank/coder/reward resident).
|
|
ERP_GPU_MEM_UTIL=0.35
|
|
ERP_MAX_MODEL_LEN=32768
|
|
ERP_MAX_NUM_SEQS=8
|
|
API_KEY=
|