Files
esh-pfi-infrastructure/persistent-memory.d/2026-08-12-erp-dual-seat-overhaul.md
T

5.9 KiB
Raw Blame History

eRP dual-seat overhaul — MeroMero-v2 + Dark-Scarlett, NVFP4A16 @ 256K on ana-ml2

[2026-08-12] Replaced the two legacy char-rp seats with home-quantized NVFP4A16 vLLM seats. Operator-driven, end to end this session.

What landed

Seat (LiteLLM alias) Model Role GPU Context
char-rp (:8016) G4-MeroMero-v2-31B (Gemma-4) non-thinking PROSE, multimodal (vision) GPU0 256K @ 2.07× (util 0.52)
char-rp-reasoning (:8018) Dark-Scarlett-v1.0-27B (Qwen3.6) THINKING (default) GPU1 256K @ 1.62× (util 0.44)
  • Both NVFP4A16 weight-only (llm-compressor, compressed-tensors), --kv-cache-dtype fp8.
  • Replace: char-rp-gguf (Magidonia-24B GGUF/llama.cpp, :8016) + heretic2-charrp-reasoning (DavidAU Qwen3.6-27B-Heretic2 modelopt NVFP4+MTP, :8018). Old stacks/containers stopped + retained for rollback.
  • Compose-ified: stacks/meromero-charrp + stacks/darkscarlett-charrp-reasoning (ana-ml2 /opt/docker/compose/, mirrored to eshpfi, commit f08b6cb) → survive reboot.
  • Research that drove picks: docs/pfi/erp-thinking-finetunes-2026.md (from the gecko-65 Booth).

Load-bearing lessons (the whole point of this file)

  1. Load via the ConditionalGeneration WRAPPER class, never AutoModelForCausalLM. For a multimodal-capable base (Gemma-4, Qwen3.6), AutoModelForCausalLM.from_pretrained + save_pretrained writes a FLAT text config (Qwen3_5TextConfig, model.layers.*) that both vLLM AND SGLang reject (SGLang: "Qwen3_5ForCausalLM has no SGLang implementation"; vLLM wants Qwen3_5ForConditionalGeneration). Loading via Qwen3_5ForConditionalGeneration / Gemma4ForConditionalGeneration keeps the wrapper config they accept. This was the DS blocker — re-quant via the wrapper fixed it (Dark-Scarlett-...-NVFP4A16-wrapper).
  2. NVFP4A16 is weight-only → DATA-FREE. llm-compressor infers DataFreePipeline; calibration data is unused (only matters for W4A4 activation quant). W4A16 chosen per NVIDIA's sm_120 long-context guidance (W4A4 KLD 2-4× worse past ~10k ctx).
  3. Load on CPU (device_map=None) so llm-compressor onloads one layer at a time. device_map= "auto" packs the whole model onto the GPU and OOMs when the card isn't fully free.
  4. Both models are KV-EFFICIENT — the "dense = KV-hungry" worry was WRONG. MeroMero (Gemma-4) uses sliding-window attention (most layers cache only a bounded window); DS (Qwen3.6) uses hybrid GatedDeltaNet linear-attention (3:1 linear:full, linear layers carry no KV). Both hit full native 256K easily. (MeroMero KV pool ~542K tokens at util 0.52.)
  5. MeroMero vision reconstruction. The finetune ships processor_config.json (image_processor inline, Gemma4ImageProcessor) but NOT preprocessor_config.json — the old-format file vLLM's feature-extractor loader wants. Even google/gemma-4-31B-it (ungated!) ships only processor_config.json. FIX: extract the image_processor section → write preprocessor_config.json verbatim, serve WITHOUT --language-model-only. Verified (model correctly ID'd a red circle). Audio is config-declared but WEIGHTLESS (0 audio tensors).
  6. GPU placement. Match the KV-heavier model to the roomier GPU. GPU0 (gen neighbor, ~54GB free) > GPU1 (utility cluster, ~45GB free). Swapped MeroMero→GPU0, DS→GPU1. Pins via compose deploy.resources.reservations.devices.

Dead ends (tried + abandoned)

  • DS via llm-compressor AutoModelForCausalLM → flat config vLLM/SGLang reject. → wrapper class.
  • DS via NVIDIA ModelOpt → modelopt↔transformers version deadlock: current transformers supports qwen3_5 but crashes modelopt's sparse-moe plugin (issubclass() on a non-class); modelopt 0.43.0 pulls an old transformers that can't load qwen3_5 at all. Abandoned.
  • DS via SGLangQwen3_5ForCausalLM has no SGLang implementation. Abandoned, but it REVEALED that both engines need the wrapper (→ the fix in lesson 1).
  • device_map="auto" for the quant → CUDA OOM in the weight observer. → device_map=None.

granite retired + gateway repoint

  • vllm-granite (granite-4.1-8b, fleet summarizer, GPU1) docker stopped (reversible) to reclaim ~13.6GB GPU1 for RP context.
  • LiteLLM (ana-docker:/opt/docker/conf/litellm/config.yaml, backed up .bak-pre-granite-down-*): granite-4.1-8b alias RETIRED — commented out, now 404s cleanly (the * wildcard→llama-swap was decommissioned 2026-06-20, so no fallthrough). summarizer + classifier REPOINTED to gen (hosted_vllm/qwen3.6-35b-a3b-heretic @ :8015, enable_thinking:false) — both verified. ⚠ This LiteLLM change is server-only / not version-controlled (a follow-up).

MTP — deferred

DS's MTP heads were dropped by the CausalLM loader; deferred, not restored (spec-decode is net-negative at RP temps: ~38-52% accept at temp 0.8-1.25, below vLLM's 0.5 cutoff). The splice-back path (splice_mtp.py in the heretic2 work dir) exists if ever wanted. MeroMero (Gemma-4) has no MTP by architecture.

On-disk / where things live

  • Quant pipelines: ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/ + /tank/aimodels/darkscarlett-nvfp4-work/ (scripts, BF16 source, NVFP4 outputs).
  • Compose stacks: ana-ml2:/opt/docker/compose/{meromero-charrp,darkscarlett-charrp-reasoning}/.
  • Gateway aliases (unchanged, port-based): char-rp→:8016, char-rp-reasoning→:8018. (char-rp was also fixed from the stale magidonia-24b-v4.3 backend model name → char-rp.)

Open follow-ups

  1. LiteLLM granite/repoint change NOT version-controlled (server + backup only).
  2. eshpfi unpushed (many commits this session incl. f08b6cb, 7bd7375, 398b58a).
  3. MTP deferred (see above).
  4. DS thinks verbosely (~13:1 reasoning:content) — eval item; consumers need generous max_tokens.
  5. MeroMero full 256K needs util 0.55 (GPU0 ~1.8GB free, tight); ran at 0.52 for headroom (~4.6GB).