ca3c984f93
Benchmarked selene against gen on selene's own job: 24 designed judge items with checkable ground truth, pairwise + absolute modes, 3 repeats, on BOTH a neutral JSON prompt and Selene's native Atla template. 288 calls, all free local. neutral JSON selene 20/24 (83%) gen 23/24 (96%) native Atla selene 21/24 (88%) gen 22/24 (92%) gen won on both templates and selene's BEST sat below gen's WORST. Selene was given its own fine-tuned template as a fairness check; it gained one point, not the three it needed. Decisive defect: selene cannot emit "tie" -- 0/2 on both templates, forcing a winner on every equivalent pair. For eval work that is the case that matters most. gen returned tie correctly on the JSON template. Selene also compressed the 1-5 scale (clustered at 2s and 4s) where gen used it fully. Selene's only win was ~3x latency, unexercised at ~60 calls/day with zero queueing. TWO NAMES, TWO DIFFERENT TREATMENTS, deliberately: - chat-judge -> repointed to gen. It is a ROLE alias and ADR-0012 says consumers bind the capability, not a concrete model. Sampler profile copied from image-judge (temp 0, top_p 1.0, top_k 1, thinking off) so the served config matches the benchmarked condition. - selene-1-mini-8b -> REMOVED. It 404s. It was NOT aliased to gen. A served-name is a contract about what the model IS; answering it with a different model hides a material change behind a stable string. Operator ruling: "never repoint a named model at a different model's endpoint -- that is intentionally misleading." Verified: the gateway now returns HTTP 400 "Invalid model name" for it. Reclaimed 17.2 GiB on ana-ml2 GPU 1 (free 1,818 -> 19,450 MiB) on a card that had under 2 GiB of headroom. gen already runs on GPU 0, so the judge role moved onto an existing seat rather than allocating anything new. Canonical litellm config synced from the host; ana-ml2 README and recommended-model-settings updated. compose.yaml kept for reference, not deployed.
602 lines
21 KiB
Markdown
602 lines
21 KiB
Markdown
# Recommended Model Inference Settings — Reference Document
|
||
|
||
> **Source:** AIPA Knowledge Base — compiled from 5 KB reference documents.
|
||
> **Status:** Reference document (KB-derived; freshness not actively maintained — verify against current upstream model docs before relying on a setting for production work).
|
||
> **Purpose:** Canonical reference for llama-server / llama.cpp inference parameters across all model families with KB-documented settings.
|
||
|
||
---
|
||
|
||
## Table of Contents
|
||
|
||
1. [NVIDIA Nemotron 3 Super (120B-A12B)](#1-nvidia-nemotron-3-super-120b-a12b)
|
||
2. [NVIDIA Nemotron 3 Nano (4B / 30B-A3B)](#2-nvidia-nemotron-3-nano-4b--30b-a3b)
|
||
3. [Qwen 3.5 Family (0.8B – 397B-A17B)](#3-qwen-35-family-08b--397b-a17b)
|
||
4. [Qwen3-Coder-Next (80B MoE)](#4-qwen3-coder-next-80b-moe)
|
||
5. [Google Gemma 4 Family (E2B – 31B)](#5-google-gemma-4-family-e2b--31b)
|
||
6. [Quick Reference Cards](#6-quick-reference-cards)
|
||
7. [Critical Warnings by Model](#7-critical-warnings-by-model)
|
||
8. [Models Without KB Settings](#8-models-without-kb-settings)
|
||
9. [PFI LiteLLM Gateway — Deployed Sampling Defaults](#9-pfi-litellm-gateway--deployed-sampling-defaults)
|
||
|
||
---
|
||
|
||
## 1. NVIDIA Nemotron 3 Super (120B-A12B)
|
||
|
||
### Model Overview
|
||
|
||
| Property | Value |
|
||
|---|---|
|
||
| Architecture | MoE — 120B total, **12B active parameters** |
|
||
| Max Context | **1,048,576** (1M tokens) |
|
||
| Recommended Starting Context | **16K or 32K** — increase gradually |
|
||
| Reasoning Tokens | `<think)>` (ID 12), `</think)>` (ID 13) |
|
||
| Positional Embeddings | **NoPE** — YaRN NOT needed |
|
||
| Best For | Multi-agent AI, high-efficiency reasoning, coding, math |
|
||
| Performance Tier | ~GPT-5.2 / Claude Opus 4.5 level |
|
||
|
||
### Inference Parameters
|
||
|
||
| Parameter | General Chat / Instruction | Tool Calling |
|
||
|---|---|---|
|
||
| `temperature` | **1.0** | **0.6** |
|
||
| `top_p` | **1.0** | **0.95** |
|
||
| `min_p` | **0.01** | **0.01** |
|
||
|
||
### Additional Settings
|
||
|
||
| Setting | Value | Notes |
|
||
|---|---|---|
|
||
| `--seed` | **3407** | Reproducibility |
|
||
| `--prio` | **2** or **3** | Priority scheduling |
|
||
| `--special` | Required | To see reasoning tokens |
|
||
| `--verbose-prompt` | Required | To see prepended `<think)>` tokens |
|
||
| `max_new_tokens` | 32,768 – 262,144 | Up to 1M |
|
||
|
||
### Quantization & Memory
|
||
|
||
| Precision | Memory Required |
|
||
|---|---|
|
||
| UD-Q2_K_XL (2-bit) | ~32–36 GB |
|
||
| **UD-Q4_K_XL (4-bit)** | **~64–72 GB** |
|
||
| 8-bit | ~128 GB |
|
||
| BF16 | ~240 GB |
|
||
|
||
### Example Command
|
||
|
||
```bash
|
||
./llama.cpp/llama-server \
|
||
--model Nemotron-3-Super-UD-Q4_K_XL.gguf \
|
||
--ctx-size 16384 \
|
||
--temp 1.0 --top-p 1.0 --min-p 0.01 \
|
||
--seed 3407 --special \
|
||
--flash-attn on \
|
||
--port 8001
|
||
```
|
||
|
||
---
|
||
|
||
## 2. NVIDIA Nemotron 3 Nano (4B / 30B-A3B)
|
||
|
||
### Model Variants
|
||
|
||
| Variant | Architecture | Context | Active Params | Best Fit |
|
||
|---|---|---|---|---|
|
||
| **Nano-4B** | Dense | 128K | 4B | Lightweight coding, math, agentic tasks |
|
||
| **Nano-30B-A3B** | MoE | 128K | 3B active | Best performance/size on 24GB devices |
|
||
|
||
### Inference Parameters (Both Variants)
|
||
|
||
| Parameter | General Chat / Instruction | Tool Calling |
|
||
|---|---|---|
|
||
| `temperature` | **1.0** | **0.6** |
|
||
| `top_p` | **1.0** | **0.95** |
|
||
| `min_p` | **0.01** | **0.01** |
|
||
|
||
### Additional Settings
|
||
|
||
| Setting | Value | Notes |
|
||
|---|---|---|
|
||
| `--seed` | **3407** | Reproducibility |
|
||
| `--special` | Not required | Standard chat template for Nano variants |
|
||
|
||
### Hardware & Quantization
|
||
|
||
#### Nano-4B
|
||
|
||
| Precision | Memory |
|
||
|---|---|
|
||
| Q8_0 (8-bit, recommended) | ~3 GB |
|
||
| 4-bit | ~5 GB |
|
||
|
||
#### Nano-30B-A3B
|
||
|
||
| Precision | Memory |
|
||
|---|---|
|
||
| **UD-Q4_K_XL (4-bit, recommended)** | **~24 GB** |
|
||
| 8-bit | ~36 GB |
|
||
|
||
### Example Commands
|
||
|
||
```bash
|
||
# Nano-4B (8-bit)
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/Nemotron-3-Nano-4B-GGUF:Q8_0 \
|
||
--ctx-size 16384 \
|
||
--temp 1.0 --top-p 1.0 --min-p 0.01 \
|
||
--seed 3407 --flash-attn on --port 8001
|
||
|
||
# Nano-30B-A3B (4-bit)
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/Nemotron-3-Nano-30B-A3B-GGUF:UD-Q4_K_XL \
|
||
--ctx-size 16384 \
|
||
--temp 1.0 --top-p 1.0 --min-p 0.01 \
|
||
--seed 3407 --flash-attn on --port 8001
|
||
```
|
||
|
||
---
|
||
|
||
## 3. Qwen 3.5 Family (0.8B – 397B-A17B)
|
||
|
||
### Model Variants
|
||
|
||
| Variant | Architecture | Context | Languages | Best Fit |
|
||
|---|---|---|---|---|
|
||
| **0.8B** | Dense | 256K | 201 | Smallest edge inference |
|
||
| **2B** | Dense | 256K | 201 | Small device inference |
|
||
| **4B** | Dense | 256K | 201 | Lightweight local use |
|
||
| **9B** | Dense | 256K | 201 | Capable small model |
|
||
| **27B** | Dense | 256K | 201 | Slightly more accurate than 35B-A3B; fits 18GB |
|
||
| **35B-A3B** | MoE (3B active) | 256K | 201 | Best speed/quality tradeoff; fits 22GB |
|
||
| **122B-A10B** | MoE (10B active) | 256K | 201 | High quality; needs ~70GB (4-bit) |
|
||
| **397B-A17B** | MoE (17B active) | 256K (extendable to 1M via YaRN) | 201 | Top-tier performance |
|
||
|
||
### 27B vs 35B-A3B Decision
|
||
|
||
- **27B** — Choose for slightly more accurate results when you can't fit a larger model.
|
||
- **35B-A3B** — Choose for much faster inference. MoE with only 3B active parameters.
|
||
|
||
### Hardware Requirements
|
||
|
||
| Variant | 3-bit | 4-bit | 6-bit | 8-bit | BF16 |
|
||
|---|---|---|---|---|---|
|
||
| **0.8B / 2B** | 3 GB | 3.5 GB | 5 GB | 7.5 GB | 9 GB |
|
||
| **4B** | 4.5 GB | 5.5 GB | 7 GB | 10 GB | 14 GB |
|
||
| **9B** | 5.5 GB | 6.5 GB | 9 GB | 13 GB | 19 GB |
|
||
| **27B** | 14 GB | 17 GB | 24 GB | 30 GB | 54 GB |
|
||
| **35B-A3B** | 17 GB | 22 GB | 30 GB | 38 GB | 70 GB |
|
||
| **122B-A10B** | 60 GB | 70 GB | 106 GB | 132 GB | 245 GB |
|
||
| **397B-A17B** | 180 GB | 214 GB | 340 GB | 512 GB | 810 GB |
|
||
|
||
### Inference Parameters
|
||
|
||
#### Thinking Mode
|
||
|
||
| Parameter | General Tasks | Precise Coding (e.g. WebDev) |
|
||
|---|---|---|
|
||
| `temperature` | **1.0** | **0.6** |
|
||
| `top_p` | **0.95** | **0.95** |
|
||
| `top_k` | **20** | **20** |
|
||
| `min_p` | **0.0** | **0.0** |
|
||
| `presence_penalty` | **1.5** | **0.0** |
|
||
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
|
||
|
||
#### Non-Thinking (Instruct) Mode
|
||
|
||
| Parameter | General Tasks | Reasoning Tasks |
|
||
|---|---|---|
|
||
| `temperature` | **0.7** | **1.0** |
|
||
| `top_p` | **0.8** | **0.95** |
|
||
| `top_k` | **20** | **20** |
|
||
| `min_p` | **0.0** | **0.0** |
|
||
| `presence_penalty` | **1.5** | **1.5** |
|
||
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
|
||
|
||
### Thinking Mode Control
|
||
|
||
Enable thinking:
|
||
```bash
|
||
--chat-template-kwargs '{"enable_thinking":true}'
|
||
```
|
||
|
||
Disable thinking:
|
||
```bash
|
||
--chat-template-kwargs '{"enable_thinking":false}'
|
||
```
|
||
|
||
#### Default Thinking Behavior by Variant
|
||
|
||
| Variant | Thinking Default |
|
||
|---|---|
|
||
| **0.8B, 2B, 4B, 9B** (Small) | **Disabled** — must explicitly enable |
|
||
| **27B, 35B-A3B, 122B-A10B, 397B-A17B** | **Enabled** — must explicitly disable if unwanted |
|
||
|
||
### Context Settings
|
||
|
||
| Setting | Value |
|
||
|---|---|
|
||
| Max context window | **262,144** (256K) |
|
||
| Context extension | Up to **1M** via YaRN |
|
||
| Recommended starting context | **16,384** (16K) for responsiveness |
|
||
| Adequate output length | **32,768** tokens |
|
||
|
||
### Quantization Notes
|
||
|
||
- All GGUFs use **Unsloth Dynamic 2.0** quantization — important layers upcasted to 8 or 16-bit even in 4-bit.
|
||
- Recommended starting point: **Dynamic 4-bit** (`UD-Q4_K_XL`).
|
||
- Minimum recommended: **Dynamic 2-bit** (`UD-Q2_K_XL`).
|
||
|
||
### Example Commands
|
||
|
||
```bash
|
||
# 35B-A3B — Thinking Mode (General)
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL \
|
||
--ctx-size 16384 \
|
||
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
|
||
--chat-template-kwargs '{"enable_thinking":true}' \
|
||
--flash-attn on --port 8001
|
||
|
||
# 9B — Thinking Enabled (small models default to disabled)
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/Qwen3.5-9B-GGUF:UD-Q4_K_XL \
|
||
--ctx-size 16384 \
|
||
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
|
||
--chat-template-kwargs '{"enable_thinking":true}' \
|
||
--flash-attn on --port 8001
|
||
```
|
||
|
||
---
|
||
|
||
## 4. Qwen3-Coder-Next (80B MoE)
|
||
|
||
### Model Overview
|
||
|
||
| Property | Value |
|
||
|---|---|
|
||
| Architecture | MoE — 80B total, **3B active parameters** |
|
||
| Max Context | **262,144** (256K) |
|
||
| Recommended Context | **32,768** for less memory use |
|
||
| Thinking Mode | **Non-reasoning only** — no `<think)>` blocks |
|
||
| Best For | Fast agentic coding, long-horizon reasoning, complex tool use |
|
||
| Performance Tier | Comparable to models with 10–20× more active parameters |
|
||
|
||
### Inference Parameters
|
||
|
||
| Parameter | Value | Notes |
|
||
|---|---|---|
|
||
| `temperature` | **1.0** | |
|
||
| `top_p` | **0.95** | |
|
||
| `top_k` | **40** | Note: higher than Qwen3.5 general |
|
||
| `min_p` | **0.01** | llama.cpp default is 0.05 — override to 0.01 |
|
||
| `repetition_penalty` | **1.0** (disabled) | Only increase if you see looping |
|
||
|
||
### Quick Reference One-Liner
|
||
|
||
```
|
||
temperature=1.0, top_p=0.95, top_k=40, min_p=0.01, repetition_penalty=1.0
|
||
```
|
||
|
||
### Hardware Requirements
|
||
|
||
| Precision | Memory Required |
|
||
|---|---|
|
||
| 3-bit (UD-IQ3_XXS) | ~34 GB |
|
||
| **4-bit (UD-Q4_K_XL)** | **~46 GB** |
|
||
| 8-bit | ~85 GB |
|
||
| BF16 | ~160 GB |
|
||
|
||
### Key Differences from Qwen3.5 General Models
|
||
|
||
- **No thinking mode** — this is a non-reasoning model; ultra-quick code responses
|
||
- **Higher `top_k`** (40 vs. 20) — broader sampling for creative code generation
|
||
- `enable_thinking` flag is not applicable
|
||
|
||
### Example Commands
|
||
|
||
```bash
|
||
# llama-server deployment
|
||
./llama.cpp/llama-server \
|
||
--model Qwen3-Coder-Next-UD-Q4_K_XL.gguf \
|
||
--alias "unsloth/Qwen3-Coder-Next" \
|
||
--seed 3407 \
|
||
--temp 1.0 --top-p 0.95 --min-p 0.01 --top-k 40 \
|
||
--ctx-size 32768 \
|
||
--flash-attn on --port 8001
|
||
```
|
||
|
||
### vLLM FP8 Dynamic (GPU Premium)
|
||
|
||
```bash
|
||
CUDA_VISIBLE_DEVICES='0,1,2,3' vllm serve unsloth/Qwen3-Coder-Next-FP8-Dynamic \
|
||
--served-model-name unsloth/Qwen3-Coder-Next \
|
||
--tensor-parallel-size 4 \
|
||
--tool-call-parser qwen3_coder \
|
||
--enable-auto-tool-choice \
|
||
--dtype bfloat16 --seed 3407 \
|
||
--max-model-len 200000 \
|
||
--gpu-memory-utilization 0.93 \
|
||
--port 8001
|
||
```
|
||
|
||
---
|
||
|
||
## 5. Google Gemma 4 Family (E2B – 31B)
|
||
|
||
### Model Variants
|
||
|
||
| Variant | Architecture | Context | Modalities | Best Fit |
|
||
|---|---|---|---|---|
|
||
| **E2B** | Dense + PLE | 128K | Text, Image, Audio | Phone / edge, ASR, speech translation |
|
||
| **E4B** | Dense + PLE | 128K | Text, Image, Audio | Laptops, fast local multimodal |
|
||
| **26B-A4B** | MoE (4B active) | 256K | Text, Image | Best speed/quality tradeoff |
|
||
| **31B** | Dense | 256K | Text, Image | Strongest performance |
|
||
|
||
### Inference Parameters (All Variants)
|
||
|
||
These are Google's default Gemma 4 parameters:
|
||
|
||
| Parameter | Value |
|
||
|---|---|
|
||
| `temperature` | **1.0** |
|
||
| `top_p` | **0.95** |
|
||
| `top_k` | **64** |
|
||
| `repetition_penalty` | **1.0** (disabled — only increase if looping) |
|
||
| End-of-sentence token | `<turn|>` |
|
||
|
||
### Context Length
|
||
|
||
- **E2B / E4B**: max **128K**
|
||
- **26B-A4B / 31B**: max **256K**
|
||
- **Practical tip**: Start with **32K** for responsiveness, then increase as needed.
|
||
|
||
### Hardware Requirements
|
||
|
||
| Variant | 4-bit | 8-bit | BF16/FP16 |
|
||
|---|---|---|---|
|
||
| **E2B** | 4 GB | 5–8 GB | 10 GB |
|
||
| **E4B** | 5.5–6 GB | 9–12 GB | 16 GB |
|
||
| **26B-A4B** | 16–18 GB | 28–30 GB | 52 GB |
|
||
| **31B** | 17–20 GB | 34–38 GB | 62 GB |
|
||
|
||
### Quantization Recommendations
|
||
|
||
- **E2B / E4B** (small): prefer **Q8_0** (8-bit) for quality.
|
||
- **26B-A4B / 31B** (large): prefer **UD-Q4_K_XL** (Dynamic 4-bit) as starting point.
|
||
|
||
### 26B-A4B vs 31B Decision
|
||
|
||
- **26B-A4B** — Choose when RAM is limited. MoE with 4B active params = faster. Slight quality tradeoff.
|
||
- **31B** — Choose when you have ≥20 GB (4-bit) and want maximum quality. Slower inference.
|
||
|
||
### Thinking Mode
|
||
|
||
**Enable thinking** — add to system prompt:
|
||
```
|
||
<|think|>
|
||
You are a careful coding assistant. Explain your answer clearly.
|
||
```
|
||
|
||
Model outputs:
|
||
```
|
||
<|channel>thought
|
||
[internal reasoning]
|
||
<channel|>
|
||
[final answer]
|
||
```
|
||
|
||
**Disable thinking** via llama-server:
|
||
```bash
|
||
--chat-template-kwargs '{"enable_thinking":false}'
|
||
```
|
||
|
||
> **Multi-turn rule:** Only keep the final visible answer in chat history. Do NOT feed prior thought blocks back.
|
||
|
||
### Multimodal Settings
|
||
|
||
- **Images/Audio** go **before text** in prompts.
|
||
- **Video**: pass frames first, then instruction.
|
||
- **Audio** is only on E2B and E4B. Max audio: 30s. Max video: 60s (1 fps).
|
||
|
||
#### Visual Token Budgets
|
||
|
||
| Budget | Use Case |
|
||
|---|---|
|
||
| 70 / 140 | Classification, captioning, fast video |
|
||
| 280 / 560 | General multimodal chat, charts, UI |
|
||
| 1120 | OCR, document parsing, handwriting |
|
||
|
||
### Example Commands
|
||
|
||
```bash
|
||
# 26B-A4B (Dynamic 4-bit)
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \
|
||
--temp 1.0 --top-p 0.95 --top-k 64 \
|
||
--ctx-size 32768 --flash-attn on --port 8001
|
||
|
||
# 31B (Dynamic 4-bit) with thinking
|
||
./llama.cpp/llama-server \
|
||
-hf unsloth/gemma-4-31B-it-GGUF:UD-Q4_K_XL \
|
||
--temp 1.0 --top-p 0.95 --top-k 64 \
|
||
--ctx-size 32768 \
|
||
--chat-template-kwargs '{"enable_thinking":true}' \
|
||
--flash-attn on --port 8001
|
||
```
|
||
|
||
---
|
||
|
||
## 6. Quick Reference Cards
|
||
|
||
### Temperature by Model & Use Case
|
||
|
||
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|
||
|---|---|---|---|---|
|
||
| **Nemotron 3 Super** | 1.0 | — | 0.6 | — |
|
||
| **Nemotron 3 Nano** | 1.0 | — | 0.6 | — |
|
||
| **Qwen 3.5** (thinking) | 1.0 | 0.6 | — | — |
|
||
| **Qwen 3.5** (non-thinking) | — | — | — | 0.7 (general) / 1.0 (reasoning) |
|
||
| **Qwen3-Coder-Next** | 1.0 | 1.0 | — | — |
|
||
| **Gemma 4** | 1.0 | 1.0 | — | — |
|
||
|
||
### Top-P by Model & Use Case
|
||
|
||
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|
||
|---|---|---|---|---|
|
||
| **Nemotron 3 Super** | 1.0 | — | 0.95 | — |
|
||
| **Nemotron 3 Nano** | 1.0 | — | 0.95 | — |
|
||
| **Qwen 3.5** (thinking) | 0.95 | 0.95 | — | — |
|
||
| **Qwen 3.5** (non-thinking) | — | — | — | 0.8 (general) / 0.95 (reasoning) |
|
||
| **Qwen3-Coder-Next** | 0.95 | 0.95 | — | — |
|
||
| **Gemma 4** | 0.95 | 0.95 | — | — |
|
||
|
||
### Top-K by Model
|
||
|
||
| Model Family | Top-K |
|
||
|---|---|
|
||
| **Nemotron 3** | Default (not specified) |
|
||
| **Qwen 3.5** | **20** |
|
||
| **Qwen3-Coder-Next** | **40** |
|
||
| **Gemma 4** | **64** |
|
||
|
||
### Min-P by Model
|
||
|
||
| Model Family | Min-P |
|
||
|---|---|
|
||
| **Nemotron 3 (all)** | **0.01** |
|
||
| **Qwen 3.5** | **0.0** |
|
||
| **Qwen3-Coder-Next** | **0.01** |
|
||
| **Gemma 4** | Default (not specified) |
|
||
|
||
### Presence Penalty by Model
|
||
|
||
| Model Family | General | Coding | Reasoning |
|
||
|---|---|---|---|
|
||
| **Nemotron 3** | Default | Default | Default |
|
||
| **Qwen 3.5** (thinking) | **1.5** | **0.0** | **1.5** |
|
||
| **Qwen 3.5** (non-thinking) | **1.5** | — | **1.5** |
|
||
| **Qwen3-Coder-Next** | Default | Default | Default |
|
||
| **Gemma 4** | Default | Default | Default |
|
||
|
||
### Repetition Penalty by Model
|
||
|
||
| Model Family | Value | Notes |
|
||
|---|---|---|
|
||
| **Nemotron 3** | Default | — |
|
||
| **Qwen 3.5** | **1.0** (disabled) | — |
|
||
| **Qwen3-Coder-Next** | **1.0** (disabled) | Only increase if looping |
|
||
| **Gemma 4** | **1.0** (disabled) | Only increase if looping |
|
||
|
||
### Seed Values
|
||
|
||
| Model Family | Recommended Seed |
|
||
|---|---|
|
||
| **Nemotron 3 (all)** | **3407** |
|
||
| **Qwen 3.5** | 3407 (optional) |
|
||
| **Qwen3-Coder-Next** | **3407** |
|
||
| **Gemma 4** | Default |
|
||
|
||
---
|
||
|
||
## 7. Critical Warnings by Model
|
||
|
||
### Nemotron 3 Super
|
||
- ⚠️ Do NOT attempt 1M context on first run — increase gradually from 16K/32K.
|
||
- ⚠️ Setting context to 1M may trigger **CUDA OOM and crash**.
|
||
- ⚠️ Router-layer fine-tuning **disabled by default** in Unsloth for MoE models.
|
||
|
||
### Qwen 3.5
|
||
- ⚠️ **No Qwen3.5 GGUF works in Ollama** due to separate mmproj vision files. Use llama.cpp-compatible backends.
|
||
- ⚠️ `presence_penalty` above 0.0 may cause **slight performance decrease**.
|
||
- ⚠️ If getting gibberish, check context length or add `--cache-type-k bf16 --cache-type-v bf16`.
|
||
- ⚠️ Redownload older GGUFs — all updated with improved quantization and tool-calling template fixes.
|
||
|
||
### Qwen3-Coder-Next
|
||
- ⚠️ **Update llama.cpp** — a previous bug in `vectorized key_gdiff` caused looping/output issues.
|
||
- ⚠️ **No Ollama support** — use llama.cpp-compatible backends.
|
||
- ⚠️ If context length too low, may see `exceeds the available context size` errors.
|
||
- ⚠️ Tool-calling improved after llama.cpp parsing fixes (Feb 19 update) — use recent version.
|
||
|
||
### Gemma 4
|
||
- ⚠️ **Do NOT use CUDA 13.2 runtime** for any GGUF — causes poor outputs.
|
||
- ⚠️ Use `llama-server` (not `llama-cli`) for thinking control — more reliable.
|
||
- ⚠️ Multi-turn: **only keep the final visible answer** in chat history. Do NOT feed prior thought blocks back.
|
||
|
||
---
|
||
|
||
## 8. Models Without KB Settings
|
||
|
||
The following model families are deployed in the Infrastructure-PFI environment but **do not have KB-documented inference parameters**. Settings for these models use general best practices or vendor defaults:
|
||
|
||
| Model | Notes |
|
||
|---|---|
|
||
| **DeepSeek R1 0528** | No KB doc — use general MoE defaults |
|
||
| **Mistral Small 3.1** | No KB doc — use vendor defaults |
|
||
| **GLM-4.7-Flash** | No KB doc — use vendor defaults |
|
||
| **GLM Steam 106B-A12B** | No KB doc — use general MoE defaults |
|
||
| **Granite 4.0 Micro** | No KB doc — use IBM defaults |
|
||
| **Kimi K2** | No KB doc — use general MoE defaults |
|
||
| **Skyfall R1 31B v4a** | No KB doc — use general defaults |
|
||
| **Hermes 4 14B** | No KB doc (download pending) — use vendor defaults |
|
||
|
||
---
|
||
|
||
## 9. PFI LiteLLM Gateway — Deployed Sampling Defaults
|
||
|
||
> **Live as of 2026-06-27** on the PFI gateway (`ana-docker:4000`; canonical config
|
||
> `eshpfi-management/stacks/litellm/conf/config.yaml`). Unlike §§1–8 (general vendor
|
||
> reference), this section is the **deployed reality** — keep it in sync when gateway
|
||
> sampling changes.
|
||
|
||
These are **overrideable defaults**: any caller that passes its own sampling param
|
||
wins; callers that omit one inherit the value below. (Verified — vLLM rejected an
|
||
out-of-range `presence_penalty=5.0`, proving per-request values reach the backend and
|
||
override the config default.) Values set per the `dvalin-smithy-dev` research pass
|
||
(provenance-cited in-thread, corroborated by §3 above). vLLM-only params (`top_k`,
|
||
`repetition_penalty`) ride in `extra_body` so LiteLLM's `drop_params` can't strip them.
|
||
|
||
| Gateway model(s) | temp | top_p | top_k | presence_penalty | repetition_penalty | Source |
|
||
|---|---|---|---|---|---|---|
|
||
| `granite-4.1-8b`, `summarizer`, `classifier` | **0** | — | — | — | — | IBM-canonical (temp 0 for inferencing) |
|
||
| `gen`, `summarizer-large`, `qwen-large`, `qwen3.5-122-a10b` (non-thinking) | **0.7** | 0.8 | 20 | **1.0** | — | Qwen3 non-thinking + operator anti-repetition |
|
||
| `gen-reasoning`, `qwen-large-reasoning`, `qwen3.5-122-a10b-reasoning` (thinking) | **0.6** | 0.95 | 20 | **1.0** | — | Qwen3 thinking |
|
||
| `qwen-image-bench`, `image-judge` | **0** | 1.0 | 1 | — | 1.05 | Qwen-Image-Bench judge reproducibility table |
|
||
| ~~`selene-1-mini-8b`~~ | — | — | — | — | — | **RETIRED 2026-08-23**; name 404s by design, not aliased |
|
||
| `chat-judge` | **0** | 1.0 | 1 | — | 1.05 | Repointed to `gen` 2026-08-23; deterministic judge profile copied from `image-judge`. The benchmark that selected `gen` ran at temperature 0 — match it. |
|
||
| `glm-5.1`, `glm-5.2`, `glm-5-turbo`, `glm-4.7`, `gen-frontier` | **1.0** | 0.95 | — | — | — | z.ai API defaults (5.x / 4.7 series) |
|
||
| `glm-4.5-air` | **0.6** | 0.95 | — | — | — | z.ai API default (4.5 series) |
|
||
| `qwen3-embedding`, `qwen3-reranker`, `reranker` | — | — | — | — | — | no sampling (embedding / rerank) |
|
||
|
||
**Notes:**
|
||
- **qwen "gen" family `presence_penalty: 1.0`** — operator-set anti-repetition for the
|
||
abliterated/NVFP4 Qwopus 122B-A10B. Qwen documents `presence_penalty` (0–2) as *the*
|
||
repetition lever; 1.0 is conservative (the §3 vendor general value is 1.5 — step up to
|
||
1.5 if loops persist). Do **not** use `repetition_penalty` for the Qwen3 family.
|
||
- **GLM (z.ai cloud) — only `temperature` + `top_p` are set.** z.ai's chat API schema
|
||
accepts no `top_k` / `min_p` / penalties, so they're deliberately not sent (would be
|
||
silently dropped). These temps match z.ai's own API defaults (explicit-over-implicit /
|
||
future-proofing).
|
||
- **Both `temp 0` values (granite, image-judge) are research-confirmed, not heuristic.**
|
||
Greedy is correct for constrained summ/classify (IBM) and for judge reproducibility
|
||
(Qwen judge card + LLM-as-judge practice). `temp 0.1` was explicitly evaluated and
|
||
rejected: it adds sampling noise without fixing loops, and *reduces* run-to-run score
|
||
consistency on the judge. If granite ever loops in production, fix via
|
||
`repetition_penalty` / `presence_penalty` / `max_tokens`, not a temperature floor.
|
||
- **`qwen-image-bench` / `image-judge` is arbo's hero-judge** (comfy-dev consumer) —
|
||
sampling changes there are a coordination item, not a unilateral gateway edit.
|
||
|
||
---
|
||
|
||
## KB Source Documents
|
||
|
||
| Document | Path in KB |
|
||
|---|---|
|
||
| Nemotron 3 Super Running Parameters | `reference/nemotron-3-super-running-parameters.md` |
|
||
| Nemotron 3 Nano Running Parameters | `reference/nemotron-3-nano-running-parameters.md` |
|
||
| Qwen 3.5 Running Parameters | `reference/qwen3.5-running-parameters.md` |
|
||
| Qwen3-Coder-Next Running Parameters | `reference/qwen3-coder-next-running-parameters.md` |
|
||
| Gemma 4 Running Parameters | `reference/gemma-4-running-parameters.md` |
|
||
|
||
---
|
||
|
||
*Document generated by Linus (Systems Architect) from AIPA Knowledge Base content curated by Atlas.*
|