feat(erp-seat): NVFP4A16 quant pipeline for the Gemma-4 26B-A4B MoE ERP tune + ana-ml2 GPU1 serve stack
- services/erp-seat-quant/quant_nvfp4a16_gemma4_moe.py: linearize_moe first (playbook §3.15), asserts the expert Linear count, routers/vision/audio/norms/lm_head ignored, W4A16 for RP long-session fidelity, post-steps restore processor configs + template and reset the tokenizer truncation cap (§3.14); --dry-run proves targets before GPU time - services/erp-seat-quant/run_quant_erp_v6.sh: detached container on GPU1 (vllm-llmcompressor) - stacks/erp-seat: serve recipe copied from gemma4-charrp, true served name only, port 8021
This commit is contained in:
@@ -0,0 +1,12 @@
|
||||
# erp-seat — ana-ml2 GPU1. Real .env lives on the host at /opt/docker/compose/erp-seat/.env.
|
||||
ERP_IMAGE=vllm/vllm-openai:v0.26.0
|
||||
ERP_MODEL=/tank/aimodels/erp-tune-v6-nvfp4a16
|
||||
ERP_SERVED_NAME=erp-tune-v6-nvfp4a16
|
||||
ERP_CHAT_TEMPLATE=/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja
|
||||
ERP_PORT=8021
|
||||
ERP_GPU_ID=1
|
||||
# 0.35 x 97.9 GiB = 34 GiB. GPU1 had ~47 GiB free on 2026-09-08 (scriberr/embed/rerank/coder/reward resident).
|
||||
ERP_GPU_MEM_UTIL=0.35
|
||||
ERP_MAX_MODEL_LEN=32768
|
||||
ERP_MAX_NUM_SEQS=8
|
||||
API_KEY=
|
||||
Reference in New Issue
Block a user