Replaces the ad-hoc docker-run seats with proper compose stacks on ana-ml2, mirrored here: - meromero-charrp: G4-MeroMero-v2-31B NVFP4A16, char-rp prose (non-thinking, multimodal, vision-enabled), GPU0, 256K @ ~2x. util 0.52 (leaves ~4.6GB GPU0 headroom). - darkscarlett-charrp-reasoning: Dark-Scarlett-v1.0-27B NVFP4A16 (Qwen wrapper recipe), char-rp-reasoning thinking seat, GPU1, 256K. MTP deferred (no spec-decode). Both survive reboot now. Supersede the retired char-rp-gguf + heretic2-charrp-reasoning stacks.
meromero-charrp — MeroMero-v2 char-rp prose seat (ana-ml2 GPU0)
The non-thinking, multimodal RP prose seat. Serves the LiteLLM char-rp alias.
- Model:
G4-MeroMero-v2-31B-NVFP4A16(Gemma-4-31B, home-quantized weight-only NVFP4A16). - Host/GPU: ana-ml2, GPU0 (co-located with
gen/ vllm-aeon-gen). - Port: :8016 → LiteLLM
char-rp. - Context: 256K (
--max-model-len 262144). Gemma-4 uses sliding-window attention → KV-efficient, ~2× concurrency at full context. - Vision: enabled (image + text).
preprocessor_config.jsonwas materialized from the model's ownprocessor_config.json(Gemma4ImageProcessor); audio is config-declared but weightless. - Tuning:
.env—MEROMERO_GPU_MEM_UTIL=0.52(leaves ~4.6 GB GPU0 headroom),MEROMERO_MAX_MODEL_LEN=262144,MEROMERO_GPU_ID=0.
Replaces the retired char-rp-gguf (Magidonia-24B GGUF / llama.cpp) seat. The quant pipeline
lives in ana-ml2:/tank/aimodels/meromero-v2-nvfp4-work/.
Deploy
scripts/deploy-stack.sh ana-ml2 meromero-charrp # diffs vs live, prompts y/N
# on host: cp .env.example .env; docker compose up -d