Files
esh-pfi-infrastructure/docs/pfi/recommended-model-settings.md
T
vh ca3c984f93 feat(selene): retire the seat; chat-judge -> gen, selene-1-mini-8b 404s by design
Benchmarked selene against gen on selene's own job: 24 designed judge items
with checkable ground truth, pairwise + absolute modes, 3 repeats, on BOTH a
neutral JSON prompt and Selene's native Atla template. 288 calls, all free
local.

  neutral JSON     selene 20/24 (83%)   gen 23/24 (96%)
  native Atla      selene 21/24 (88%)   gen 22/24 (92%)

gen won on both templates and selene's BEST sat below gen's WORST. Selene was
given its own fine-tuned template as a fairness check; it gained one point,
not the three it needed.

Decisive defect: selene cannot emit "tie" -- 0/2 on both templates, forcing a
winner on every equivalent pair. For eval work that is the case that matters
most. gen returned tie correctly on the JSON template. Selene also compressed
the 1-5 scale (clustered at 2s and 4s) where gen used it fully. Selene's only
win was ~3x latency, unexercised at ~60 calls/day with zero queueing.

TWO NAMES, TWO DIFFERENT TREATMENTS, deliberately:

- chat-judge -> repointed to gen. It is a ROLE alias and ADR-0012 says
  consumers bind the capability, not a concrete model. Sampler profile copied
  from image-judge (temp 0, top_p 1.0, top_k 1, thinking off) so the served
  config matches the benchmarked condition.

- selene-1-mini-8b -> REMOVED. It 404s. It was NOT aliased to gen. A
  served-name is a contract about what the model IS; answering it with a
  different model hides a material change behind a stable string. Operator
  ruling: "never repoint a named model at a different model's endpoint --
  that is intentionally misleading." Verified: the gateway now returns
  HTTP 400 "Invalid model name" for it.

Reclaimed 17.2 GiB on ana-ml2 GPU 1 (free 1,818 -> 19,450 MiB) on a card that
had under 2 GiB of headroom. gen already runs on GPU 0, so the judge role
moved onto an existing seat rather than allocating anything new.

Canonical litellm config synced from the host; ana-ml2 README and
recommended-model-settings updated. compose.yaml kept for reference, not
deployed.
2026-08-23 05:07:01 -07:00

602 lines
21 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Recommended Model Inference Settings — Reference Document
> **Source:** AIPA Knowledge Base — compiled from 5 KB reference documents.
> **Status:** Reference document (KB-derived; freshness not actively maintained — verify against current upstream model docs before relying on a setting for production work).
> **Purpose:** Canonical reference for llama-server / llama.cpp inference parameters across all model families with KB-documented settings.
---
## Table of Contents
1. [NVIDIA Nemotron 3 Super (120B-A12B)](#1-nvidia-nemotron-3-super-120b-a12b)
2. [NVIDIA Nemotron 3 Nano (4B / 30B-A3B)](#2-nvidia-nemotron-3-nano-4b--30b-a3b)
3. [Qwen 3.5 Family (0.8B 397B-A17B)](#3-qwen-35-family-08b--397b-a17b)
4. [Qwen3-Coder-Next (80B MoE)](#4-qwen3-coder-next-80b-moe)
5. [Google Gemma 4 Family (E2B 31B)](#5-google-gemma-4-family-e2b--31b)
6. [Quick Reference Cards](#6-quick-reference-cards)
7. [Critical Warnings by Model](#7-critical-warnings-by-model)
8. [Models Without KB Settings](#8-models-without-kb-settings)
9. [PFI LiteLLM Gateway — Deployed Sampling Defaults](#9-pfi-litellm-gateway--deployed-sampling-defaults)
---
## 1. NVIDIA Nemotron 3 Super (120B-A12B)
### Model Overview
| Property | Value |
|---|---|
| Architecture | MoE — 120B total, **12B active parameters** |
| Max Context | **1,048,576** (1M tokens) |
| Recommended Starting Context | **16K or 32K** — increase gradually |
| Reasoning Tokens | `<think)>` (ID 12), `</think)>` (ID 13) |
| Positional Embeddings | **NoPE** — YaRN NOT needed |
| Best For | Multi-agent AI, high-efficiency reasoning, coding, math |
| Performance Tier | ~GPT-5.2 / Claude Opus 4.5 level |
### Inference Parameters
| Parameter | General Chat / Instruction | Tool Calling |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **1.0** | **0.95** |
| `min_p` | **0.01** | **0.01** |
### Additional Settings
| Setting | Value | Notes |
|---|---|---|
| `--seed` | **3407** | Reproducibility |
| `--prio` | **2** or **3** | Priority scheduling |
| `--special` | Required | To see reasoning tokens |
| `--verbose-prompt` | Required | To see prepended `<think)>` tokens |
| `max_new_tokens` | 32,768 262,144 | Up to 1M |
### Quantization & Memory
| Precision | Memory Required |
|---|---|
| UD-Q2_K_XL (2-bit) | ~3236 GB |
| **UD-Q4_K_XL (4-bit)** | **~6472 GB** |
| 8-bit | ~128 GB |
| BF16 | ~240 GB |
### Example Command
```bash
./llama.cpp/llama-server \
--model Nemotron-3-Super-UD-Q4_K_XL.gguf \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --special \
--flash-attn on \
--port 8001
```
---
## 2. NVIDIA Nemotron 3 Nano (4B / 30B-A3B)
### Model Variants
| Variant | Architecture | Context | Active Params | Best Fit |
|---|---|---|---|---|
| **Nano-4B** | Dense | 128K | 4B | Lightweight coding, math, agentic tasks |
| **Nano-30B-A3B** | MoE | 128K | 3B active | Best performance/size on 24GB devices |
### Inference Parameters (Both Variants)
| Parameter | General Chat / Instruction | Tool Calling |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **1.0** | **0.95** |
| `min_p` | **0.01** | **0.01** |
### Additional Settings
| Setting | Value | Notes |
|---|---|---|
| `--seed` | **3407** | Reproducibility |
| `--special` | Not required | Standard chat template for Nano variants |
### Hardware & Quantization
#### Nano-4B
| Precision | Memory |
|---|---|
| Q8_0 (8-bit, recommended) | ~3 GB |
| 4-bit | ~5 GB |
#### Nano-30B-A3B
| Precision | Memory |
|---|---|
| **UD-Q4_K_XL (4-bit, recommended)** | **~24 GB** |
| 8-bit | ~36 GB |
### Example Commands
```bash
# Nano-4B (8-bit)
./llama.cpp/llama-server \
-hf unsloth/Nemotron-3-Nano-4B-GGUF:Q8_0 \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --flash-attn on --port 8001
# Nano-30B-A3B (4-bit)
./llama.cpp/llama-server \
-hf unsloth/Nemotron-3-Nano-30B-A3B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --flash-attn on --port 8001
```
---
## 3. Qwen 3.5 Family (0.8B 397B-A17B)
### Model Variants
| Variant | Architecture | Context | Languages | Best Fit |
|---|---|---|---|---|
| **0.8B** | Dense | 256K | 201 | Smallest edge inference |
| **2B** | Dense | 256K | 201 | Small device inference |
| **4B** | Dense | 256K | 201 | Lightweight local use |
| **9B** | Dense | 256K | 201 | Capable small model |
| **27B** | Dense | 256K | 201 | Slightly more accurate than 35B-A3B; fits 18GB |
| **35B-A3B** | MoE (3B active) | 256K | 201 | Best speed/quality tradeoff; fits 22GB |
| **122B-A10B** | MoE (10B active) | 256K | 201 | High quality; needs ~70GB (4-bit) |
| **397B-A17B** | MoE (17B active) | 256K (extendable to 1M via YaRN) | 201 | Top-tier performance |
### 27B vs 35B-A3B Decision
- **27B** — Choose for slightly more accurate results when you can't fit a larger model.
- **35B-A3B** — Choose for much faster inference. MoE with only 3B active parameters.
### Hardware Requirements
| Variant | 3-bit | 4-bit | 6-bit | 8-bit | BF16 |
|---|---|---|---|---|---|
| **0.8B / 2B** | 3 GB | 3.5 GB | 5 GB | 7.5 GB | 9 GB |
| **4B** | 4.5 GB | 5.5 GB | 7 GB | 10 GB | 14 GB |
| **9B** | 5.5 GB | 6.5 GB | 9 GB | 13 GB | 19 GB |
| **27B** | 14 GB | 17 GB | 24 GB | 30 GB | 54 GB |
| **35B-A3B** | 17 GB | 22 GB | 30 GB | 38 GB | 70 GB |
| **122B-A10B** | 60 GB | 70 GB | 106 GB | 132 GB | 245 GB |
| **397B-A17B** | 180 GB | 214 GB | 340 GB | 512 GB | 810 GB |
### Inference Parameters
#### Thinking Mode
| Parameter | General Tasks | Precise Coding (e.g. WebDev) |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **0.95** | **0.95** |
| `top_k` | **20** | **20** |
| `min_p` | **0.0** | **0.0** |
| `presence_penalty` | **1.5** | **0.0** |
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
#### Non-Thinking (Instruct) Mode
| Parameter | General Tasks | Reasoning Tasks |
|---|---|---|
| `temperature` | **0.7** | **1.0** |
| `top_p` | **0.8** | **0.95** |
| `top_k` | **20** | **20** |
| `min_p` | **0.0** | **0.0** |
| `presence_penalty` | **1.5** | **1.5** |
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
### Thinking Mode Control
Enable thinking:
```bash
--chat-template-kwargs '{"enable_thinking":true}'
```
Disable thinking:
```bash
--chat-template-kwargs '{"enable_thinking":false}'
```
#### Default Thinking Behavior by Variant
| Variant | Thinking Default |
|---|---|
| **0.8B, 2B, 4B, 9B** (Small) | **Disabled** — must explicitly enable |
| **27B, 35B-A3B, 122B-A10B, 397B-A17B** | **Enabled** — must explicitly disable if unwanted |
### Context Settings
| Setting | Value |
|---|---|
| Max context window | **262,144** (256K) |
| Context extension | Up to **1M** via YaRN |
| Recommended starting context | **16,384** (16K) for responsiveness |
| Adequate output length | **32,768** tokens |
### Quantization Notes
- All GGUFs use **Unsloth Dynamic 2.0** quantization — important layers upcasted to 8 or 16-bit even in 4-bit.
- Recommended starting point: **Dynamic 4-bit** (`UD-Q4_K_XL`).
- Minimum recommended: **Dynamic 2-bit** (`UD-Q2_K_XL`).
### Example Commands
```bash
# 35B-A3B — Thinking Mode (General)
./llama.cpp/llama-server \
-hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
# 9B — Thinking Enabled (small models default to disabled)
./llama.cpp/llama-server \
-hf unsloth/Qwen3.5-9B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
```
---
## 4. Qwen3-Coder-Next (80B MoE)
### Model Overview
| Property | Value |
|---|---|
| Architecture | MoE — 80B total, **3B active parameters** |
| Max Context | **262,144** (256K) |
| Recommended Context | **32,768** for less memory use |
| Thinking Mode | **Non-reasoning only** — no `<think)>` blocks |
| Best For | Fast agentic coding, long-horizon reasoning, complex tool use |
| Performance Tier | Comparable to models with 1020× more active parameters |
### Inference Parameters
| Parameter | Value | Notes |
|---|---|---|
| `temperature` | **1.0** | |
| `top_p` | **0.95** | |
| `top_k` | **40** | Note: higher than Qwen3.5 general |
| `min_p` | **0.01** | llama.cpp default is 0.05 — override to 0.01 |
| `repetition_penalty` | **1.0** (disabled) | Only increase if you see looping |
### Quick Reference One-Liner
```
temperature=1.0, top_p=0.95, top_k=40, min_p=0.01, repetition_penalty=1.0
```
### Hardware Requirements
| Precision | Memory Required |
|---|---|
| 3-bit (UD-IQ3_XXS) | ~34 GB |
| **4-bit (UD-Q4_K_XL)** | **~46 GB** |
| 8-bit | ~85 GB |
| BF16 | ~160 GB |
### Key Differences from Qwen3.5 General Models
- **No thinking mode** — this is a non-reasoning model; ultra-quick code responses
- **Higher `top_k`** (40 vs. 20) — broader sampling for creative code generation
- `enable_thinking` flag is not applicable
### Example Commands
```bash
# llama-server deployment
./llama.cpp/llama-server \
--model Qwen3-Coder-Next-UD-Q4_K_XL.gguf \
--alias "unsloth/Qwen3-Coder-Next" \
--seed 3407 \
--temp 1.0 --top-p 0.95 --min-p 0.01 --top-k 40 \
--ctx-size 32768 \
--flash-attn on --port 8001
```
### vLLM FP8 Dynamic (GPU Premium)
```bash
CUDA_VISIBLE_DEVICES='0,1,2,3' vllm serve unsloth/Qwen3-Coder-Next-FP8-Dynamic \
--served-model-name unsloth/Qwen3-Coder-Next \
--tensor-parallel-size 4 \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice \
--dtype bfloat16 --seed 3407 \
--max-model-len 200000 \
--gpu-memory-utilization 0.93 \
--port 8001
```
---
## 5. Google Gemma 4 Family (E2B 31B)
### Model Variants
| Variant | Architecture | Context | Modalities | Best Fit |
|---|---|---|---|---|
| **E2B** | Dense + PLE | 128K | Text, Image, Audio | Phone / edge, ASR, speech translation |
| **E4B** | Dense + PLE | 128K | Text, Image, Audio | Laptops, fast local multimodal |
| **26B-A4B** | MoE (4B active) | 256K | Text, Image | Best speed/quality tradeoff |
| **31B** | Dense | 256K | Text, Image | Strongest performance |
### Inference Parameters (All Variants)
These are Google's default Gemma 4 parameters:
| Parameter | Value |
|---|---|
| `temperature` | **1.0** |
| `top_p` | **0.95** |
| `top_k` | **64** |
| `repetition_penalty` | **1.0** (disabled — only increase if looping) |
| End-of-sentence token | `<turn|>` |
### Context Length
- **E2B / E4B**: max **128K**
- **26B-A4B / 31B**: max **256K**
- **Practical tip**: Start with **32K** for responsiveness, then increase as needed.
### Hardware Requirements
| Variant | 4-bit | 8-bit | BF16/FP16 |
|---|---|---|---|
| **E2B** | 4 GB | 58 GB | 10 GB |
| **E4B** | 5.56 GB | 912 GB | 16 GB |
| **26B-A4B** | 1618 GB | 2830 GB | 52 GB |
| **31B** | 1720 GB | 3438 GB | 62 GB |
### Quantization Recommendations
- **E2B / E4B** (small): prefer **Q8_0** (8-bit) for quality.
- **26B-A4B / 31B** (large): prefer **UD-Q4_K_XL** (Dynamic 4-bit) as starting point.
### 26B-A4B vs 31B Decision
- **26B-A4B** — Choose when RAM is limited. MoE with 4B active params = faster. Slight quality tradeoff.
- **31B** — Choose when you have ≥20 GB (4-bit) and want maximum quality. Slower inference.
### Thinking Mode
**Enable thinking** — add to system prompt:
```
<|think|>
You are a careful coding assistant. Explain your answer clearly.
```
Model outputs:
```
<|channel>thought
[internal reasoning]
<channel|>
[final answer]
```
**Disable thinking** via llama-server:
```bash
--chat-template-kwargs '{"enable_thinking":false}'
```
> **Multi-turn rule:** Only keep the final visible answer in chat history. Do NOT feed prior thought blocks back.
### Multimodal Settings
- **Images/Audio** go **before text** in prompts.
- **Video**: pass frames first, then instruction.
- **Audio** is only on E2B and E4B. Max audio: 30s. Max video: 60s (1 fps).
#### Visual Token Budgets
| Budget | Use Case |
|---|---|
| 70 / 140 | Classification, captioning, fast video |
| 280 / 560 | General multimodal chat, charts, UI |
| 1120 | OCR, document parsing, handwriting |
### Example Commands
```bash
# 26B-A4B (Dynamic 4-bit)
./llama.cpp/llama-server \
-hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \
--temp 1.0 --top-p 0.95 --top-k 64 \
--ctx-size 32768 --flash-attn on --port 8001
# 31B (Dynamic 4-bit) with thinking
./llama.cpp/llama-server \
-hf unsloth/gemma-4-31B-it-GGUF:UD-Q4_K_XL \
--temp 1.0 --top-p 0.95 --top-k 64 \
--ctx-size 32768 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
```
---
## 6. Quick Reference Cards
### Temperature by Model & Use Case
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|---|---|---|---|---|
| **Nemotron 3 Super** | 1.0 | — | 0.6 | — |
| **Nemotron 3 Nano** | 1.0 | — | 0.6 | — |
| **Qwen 3.5** (thinking) | 1.0 | 0.6 | — | — |
| **Qwen 3.5** (non-thinking) | — | — | — | 0.7 (general) / 1.0 (reasoning) |
| **Qwen3-Coder-Next** | 1.0 | 1.0 | — | — |
| **Gemma 4** | 1.0 | 1.0 | — | — |
### Top-P by Model & Use Case
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|---|---|---|---|---|
| **Nemotron 3 Super** | 1.0 | — | 0.95 | — |
| **Nemotron 3 Nano** | 1.0 | — | 0.95 | — |
| **Qwen 3.5** (thinking) | 0.95 | 0.95 | — | — |
| **Qwen 3.5** (non-thinking) | — | — | — | 0.8 (general) / 0.95 (reasoning) |
| **Qwen3-Coder-Next** | 0.95 | 0.95 | — | — |
| **Gemma 4** | 0.95 | 0.95 | — | — |
### Top-K by Model
| Model Family | Top-K |
|---|---|
| **Nemotron 3** | Default (not specified) |
| **Qwen 3.5** | **20** |
| **Qwen3-Coder-Next** | **40** |
| **Gemma 4** | **64** |
### Min-P by Model
| Model Family | Min-P |
|---|---|
| **Nemotron 3 (all)** | **0.01** |
| **Qwen 3.5** | **0.0** |
| **Qwen3-Coder-Next** | **0.01** |
| **Gemma 4** | Default (not specified) |
### Presence Penalty by Model
| Model Family | General | Coding | Reasoning |
|---|---|---|---|
| **Nemotron 3** | Default | Default | Default |
| **Qwen 3.5** (thinking) | **1.5** | **0.0** | **1.5** |
| **Qwen 3.5** (non-thinking) | **1.5** | — | **1.5** |
| **Qwen3-Coder-Next** | Default | Default | Default |
| **Gemma 4** | Default | Default | Default |
### Repetition Penalty by Model
| Model Family | Value | Notes |
|---|---|---|
| **Nemotron 3** | Default | — |
| **Qwen 3.5** | **1.0** (disabled) | — |
| **Qwen3-Coder-Next** | **1.0** (disabled) | Only increase if looping |
| **Gemma 4** | **1.0** (disabled) | Only increase if looping |
### Seed Values
| Model Family | Recommended Seed |
|---|---|
| **Nemotron 3 (all)** | **3407** |
| **Qwen 3.5** | 3407 (optional) |
| **Qwen3-Coder-Next** | **3407** |
| **Gemma 4** | Default |
---
## 7. Critical Warnings by Model
### Nemotron 3 Super
- ⚠️ Do NOT attempt 1M context on first run — increase gradually from 16K/32K.
- ⚠️ Setting context to 1M may trigger **CUDA OOM and crash**.
- ⚠️ Router-layer fine-tuning **disabled by default** in Unsloth for MoE models.
### Qwen 3.5
- ⚠️ **No Qwen3.5 GGUF works in Ollama** due to separate mmproj vision files. Use llama.cpp-compatible backends.
- ⚠️ `presence_penalty` above 0.0 may cause **slight performance decrease**.
- ⚠️ If getting gibberish, check context length or add `--cache-type-k bf16 --cache-type-v bf16`.
- ⚠️ Redownload older GGUFs — all updated with improved quantization and tool-calling template fixes.
### Qwen3-Coder-Next
- ⚠️ **Update llama.cpp** — a previous bug in `vectorized key_gdiff` caused looping/output issues.
- ⚠️ **No Ollama support** — use llama.cpp-compatible backends.
- ⚠️ If context length too low, may see `exceeds the available context size` errors.
- ⚠️ Tool-calling improved after llama.cpp parsing fixes (Feb 19 update) — use recent version.
### Gemma 4
- ⚠️ **Do NOT use CUDA 13.2 runtime** for any GGUF — causes poor outputs.
- ⚠️ Use `llama-server` (not `llama-cli`) for thinking control — more reliable.
- ⚠️ Multi-turn: **only keep the final visible answer** in chat history. Do NOT feed prior thought blocks back.
---
## 8. Models Without KB Settings
The following model families are deployed in the Infrastructure-PFI environment but **do not have KB-documented inference parameters**. Settings for these models use general best practices or vendor defaults:
| Model | Notes |
|---|---|
| **DeepSeek R1 0528** | No KB doc — use general MoE defaults |
| **Mistral Small 3.1** | No KB doc — use vendor defaults |
| **GLM-4.7-Flash** | No KB doc — use vendor defaults |
| **GLM Steam 106B-A12B** | No KB doc — use general MoE defaults |
| **Granite 4.0 Micro** | No KB doc — use IBM defaults |
| **Kimi K2** | No KB doc — use general MoE defaults |
| **Skyfall R1 31B v4a** | No KB doc — use general defaults |
| **Hermes 4 14B** | No KB doc (download pending) — use vendor defaults |
---
## 9. PFI LiteLLM Gateway — Deployed Sampling Defaults
> **Live as of 2026-06-27** on the PFI gateway (`ana-docker:4000`; canonical config
> `eshpfi-management/stacks/litellm/conf/config.yaml`). Unlike §§18 (general vendor
> reference), this section is the **deployed reality** — keep it in sync when gateway
> sampling changes.
These are **overrideable defaults**: any caller that passes its own sampling param
wins; callers that omit one inherit the value below. (Verified — vLLM rejected an
out-of-range `presence_penalty=5.0`, proving per-request values reach the backend and
override the config default.) Values set per the `dvalin-smithy-dev` research pass
(provenance-cited in-thread, corroborated by §3 above). vLLM-only params (`top_k`,
`repetition_penalty`) ride in `extra_body` so LiteLLM's `drop_params` can't strip them.
| Gateway model(s) | temp | top_p | top_k | presence_penalty | repetition_penalty | Source |
|---|---|---|---|---|---|---|
| `granite-4.1-8b`, `summarizer`, `classifier` | **0** | — | — | — | — | IBM-canonical (temp 0 for inferencing) |
| `gen`, `summarizer-large`, `qwen-large`, `qwen3.5-122-a10b` (non-thinking) | **0.7** | 0.8 | 20 | **1.0** | — | Qwen3 non-thinking + operator anti-repetition |
| `gen-reasoning`, `qwen-large-reasoning`, `qwen3.5-122-a10b-reasoning` (thinking) | **0.6** | 0.95 | 20 | **1.0** | — | Qwen3 thinking |
| `qwen-image-bench`, `image-judge` | **0** | 1.0 | 1 | — | 1.05 | Qwen-Image-Bench judge reproducibility table |
| ~~`selene-1-mini-8b`~~ | — | — | — | — | — | **RETIRED 2026-08-23**; name 404s by design, not aliased |
| `chat-judge` | **0** | 1.0 | 1 | — | 1.05 | Repointed to `gen` 2026-08-23; deterministic judge profile copied from `image-judge`. The benchmark that selected `gen` ran at temperature 0 — match it. |
| `glm-5.1`, `glm-5.2`, `glm-5-turbo`, `glm-4.7`, `gen-frontier` | **1.0** | 0.95 | — | — | — | z.ai API defaults (5.x / 4.7 series) |
| `glm-4.5-air` | **0.6** | 0.95 | — | — | — | z.ai API default (4.5 series) |
| `qwen3-embedding`, `qwen3-reranker`, `reranker` | — | — | — | — | — | no sampling (embedding / rerank) |
**Notes:**
- **qwen "gen" family `presence_penalty: 1.0`** — operator-set anti-repetition for the
abliterated/NVFP4 Qwopus 122B-A10B. Qwen documents `presence_penalty` (02) as *the*
repetition lever; 1.0 is conservative (the §3 vendor general value is 1.5 — step up to
1.5 if loops persist). Do **not** use `repetition_penalty` for the Qwen3 family.
- **GLM (z.ai cloud) — only `temperature` + `top_p` are set.** z.ai's chat API schema
accepts no `top_k` / `min_p` / penalties, so they're deliberately not sent (would be
silently dropped). These temps match z.ai's own API defaults (explicit-over-implicit /
future-proofing).
- **Both `temp 0` values (granite, image-judge) are research-confirmed, not heuristic.**
Greedy is correct for constrained summ/classify (IBM) and for judge reproducibility
(Qwen judge card + LLM-as-judge practice). `temp 0.1` was explicitly evaluated and
rejected: it adds sampling noise without fixing loops, and *reduces* run-to-run score
consistency on the judge. If granite ever loops in production, fix via
`repetition_penalty` / `presence_penalty` / `max_tokens`, not a temperature floor.
- **`qwen-image-bench` / `image-judge` is arbo's hero-judge** (comfy-dev consumer) —
sampling changes there are a coordination item, not a unilateral gateway edit.
---
## KB Source Documents
| Document | Path in KB |
|---|---|
| Nemotron 3 Super Running Parameters | `reference/nemotron-3-super-running-parameters.md` |
| Nemotron 3 Nano Running Parameters | `reference/nemotron-3-nano-running-parameters.md` |
| Qwen 3.5 Running Parameters | `reference/qwen3.5-running-parameters.md` |
| Qwen3-Coder-Next Running Parameters | `reference/qwen3-coder-next-running-parameters.md` |
| Gemma 4 Running Parameters | `reference/gemma-4-running-parameters.md` |
---
*Document generated by Linus (Systems Architect) from AIPA Knowledge Base content curated by Atlas.*