feat(erp-seat): run 7 quantized to NVFP4A16 and serving as the trial seat on ana-ml2

- services/erp-seat-quant/run_quant_erp_v7.sh: v6 runner retargeted; dry-run gate
  passed identically (11,725 targets, 11,520 experts = 30x128x3, routers+vision BF16)
- 49 GiB bf16 relayed gx10 -> ana-ml2 (no key path either way; nh3-dev relays),
  checksums verified against source; quant 49 -> 16 GiB, all post-steps clean
- stacks/erp-seat: .env-driven swap to erp-tune-v7-nvfp4a16, served under its TRUE
  name; homepage labels + README updated, v6 rollback path recorded
- stacks/litellm: trial -> erp-tune-v7-nvfp4a16 (config-file alias; /model/update
  refuses a config model, so this is an edit + restart)
This commit is contained in:
vh
2026-09-09 15:30:15 -07:00
parent c335c38c19
commit 6972e7ef7f
5 changed files with 60 additions and 26 deletions
+11 -4
View File
@@ -1,11 +1,18 @@
# erp-seat — ERP-tune seat on ana-ml2 (GPU1, `:8021`)
Serves the latest gated ERP LoRA merge as an **NVFP4A16** (weight-only) compressed-tensors
checkpoint so the GX10 is free to train the next run. First occupant: **run 6** —
`erp-tune-v6-nvfp4a16` = merged-run06 (jenerallee78 ARA-abliterated Gemma-4-26B-A4B-it, index
`33c59654…`, + R47 SFT r6) quantized by `services/erp-seat-quant/`.
checkpoint so the GX10 is free to train the next run.
- **True name only.** `--served-model-name erp-tune-v6-nvfp4a16`. Gateway aliases (`trial`) are
**Current occupant: run 7** — `erp-tune-v7-nvfp4a16` = merged-run07 (jenerallee78
ARA-abliterated Gemma-4-26B-A4B-it, index `33c59654…`, + R47 SFT r7 = r6 plus the opening-split
slot and its companion loss mask), quantized by `services/erp-seat-quant/` on 2026-09-09.
49 GiB bf16 → 16 GiB NVFP4A16. Runbook `docs/runbooks/gx10-run-07.md`.
*Previous: run 6 (`erp-tune-v6-nvfp4a16`, 2026-09-08). Its artifact is still on `/tank/aimodels/`
and the pre-swap host env is at `/tmp/erp-seat-env.v6.bak` on ana-ml2, so a rollback is an `.env`
flip plus `docker compose up -d`.*
- **True name only.** `--served-model-name erp-tune-v7-nvfp4a16`. Gateway aliases (`trial`) are
set in LiteLLM on the operator's word, never here (no silent substitution — the bf16 arm on the
GX10 and this NVFP4 arm are different artifacts).
- **Recipe** = `stacks/gemma4-charrp` (same arch + format, proven on this box): `gemma4` tool
+6 -6
View File
@@ -1,5 +1,5 @@
# erp-seat — the ERP-tune seat on ana-ml2 GPU1: NVFP4A16 quant of the latest gated ERP LoRA merge
# (run 6 = jenerallee78 ARA-abliterated Gemma-4-26B-A4B + R47 SFT), served under its TRUE name.
# (run 7 = jenerallee78 ARA-abliterated Gemma-4-26B-A4B + R47 SFT + the opening-split slot), served under its TRUE name.
# Routing aliases (e.g. LiteLLM `trial`) are the operator's call and live in the gateway, not here.
#
# Serve recipe copied from stacks/gemma4-charrp (same architecture + quant format, proven on this
@@ -23,11 +23,11 @@ services:
environment:
- VLLM_API_KEY=${API_KEY:-}
command:
- ${ERP_MODEL:-/tank/aimodels/erp-tune-v6-nvfp4a16}
- ${ERP_MODEL:-/tank/aimodels/erp-tune-v7-nvfp4a16}
- --quantization
- compressed-tensors
- --served-model-name
- ${ERP_SERVED_NAME:-erp-tune-v6-nvfp4a16}
- ${ERP_SERVED_NAME:-erp-tune-v7-nvfp4a16}
- --tool-call-parser
- gemma4
- --enable-auto-tool-choice
@@ -47,7 +47,7 @@ services:
# an empty turn. The flag drops the tools from the prompt so the model answers in prose.
- --exclude-tools-when-tool-choice-none
- --chat-template
- ${ERP_CHAT_TEMPLATE:-/tank/aimodels/erp-tune-v6-nvfp4a16/chat_template.jinja}
- ${ERP_CHAT_TEMPLATE:-/tank/aimodels/erp-tune-v7-nvfp4a16/chat_template.jinja}
- --max-model-len
- "${ERP_MAX_MODEL_LEN:-32768}"
- --max-num-seqs
@@ -76,9 +76,9 @@ services:
- tnet
labels:
- homepage.group=AI - Inference
- homepage.name=erp-tune-v6 (Gemma-4 26B-A4B ARA, NVFP4A16)
- homepage.name=erp-tune-v7 (Gemma-4 26B-A4B ARA, NVFP4A16)
- homepage.icon=mdi-fire
- homepage.description=ERP-seat run-6 LoRA merge on the jenerallee78 abliteration, NVFP4A16 MoE (ana-ml2 GPU1)
- homepage.description=ERP-seat run-7 LoRA merge on the jenerallee78 abliteration, NVFP4A16 MoE (ana-ml2 GPU1)
- homepage.href=http://10.250.50.54:${ERP_PORT:-8021}/docs
networks: