# Recommended Model Inference Settings — Reference Document
> **Source:** AIPA Knowledge Base — compiled from 5 KB reference documents.
> **Status:** Reference document (KB-derived; freshness not actively maintained — verify against current upstream model docs before relying on a setting for production work).
> **Purpose:** Canonical reference for llama-server / llama.cpp inference parameters across all model families with KB-documented settings.
---
## Table of Contents
1. [NVIDIA Nemotron 3 Super (120B-A12B)](#1-nvidia-nemotron-3-super-120b-a12b)
2. [NVIDIA Nemotron 3 Nano (4B / 30B-A3B)](#2-nvidia-nemotron-3-nano-4b--30b-a3b)
3. [Qwen 3.5 Family (0.8B – 397B-A17B)](#3-qwen-35-family-08b--397b-a17b)
4. [Qwen3-Coder-Next (80B MoE)](#4-qwen3-coder-next-80b-moe)
5. [Google Gemma 4 Family (E2B – 31B)](#5-google-gemma-4-family-e2b--31b)
6. [Quick Reference Cards](#6-quick-reference-cards)
7. [Critical Warnings by Model](#7-critical-warnings-by-model)
8. [Models Without KB Settings](#8-models-without-kb-settings)
9. [PFI LiteLLM Gateway — Deployed Sampling Defaults](#9-pfi-litellm-gateway--deployed-sampling-defaults)
---
## 1. NVIDIA Nemotron 3 Super (120B-A12B)
### Model Overview
| Property | Value |
|---|---|
| Architecture | MoE — 120B total, **12B active parameters** |
| Max Context | **1,048,576** (1M tokens) |
| Recommended Starting Context | **16K or 32K** — increase gradually |
| Reasoning Tokens | `` (ID 12), `` (ID 13) |
| Positional Embeddings | **NoPE** — YaRN NOT needed |
| Best For | Multi-agent AI, high-efficiency reasoning, coding, math |
| Performance Tier | ~GPT-5.2 / Claude Opus 4.5 level |
### Inference Parameters
| Parameter | General Chat / Instruction | Tool Calling |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **1.0** | **0.95** |
| `min_p` | **0.01** | **0.01** |
### Additional Settings
| Setting | Value | Notes |
|---|---|---|
| `--seed` | **3407** | Reproducibility |
| `--prio` | **2** or **3** | Priority scheduling |
| `--special` | Required | To see reasoning tokens |
| `--verbose-prompt` | Required | To see prepended `` tokens |
| `max_new_tokens` | 32,768 – 262,144 | Up to 1M |
### Quantization & Memory
| Precision | Memory Required |
|---|---|
| UD-Q2_K_XL (2-bit) | ~32–36 GB |
| **UD-Q4_K_XL (4-bit)** | **~64–72 GB** |
| 8-bit | ~128 GB |
| BF16 | ~240 GB |
### Example Command
```bash
./llama.cpp/llama-server \
--model Nemotron-3-Super-UD-Q4_K_XL.gguf \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --special \
--flash-attn on \
--port 8001
```
---
## 2. NVIDIA Nemotron 3 Nano (4B / 30B-A3B)
### Model Variants
| Variant | Architecture | Context | Active Params | Best Fit |
|---|---|---|---|---|
| **Nano-4B** | Dense | 128K | 4B | Lightweight coding, math, agentic tasks |
| **Nano-30B-A3B** | MoE | 128K | 3B active | Best performance/size on 24GB devices |
### Inference Parameters (Both Variants)
| Parameter | General Chat / Instruction | Tool Calling |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **1.0** | **0.95** |
| `min_p` | **0.01** | **0.01** |
### Additional Settings
| Setting | Value | Notes |
|---|---|---|
| `--seed` | **3407** | Reproducibility |
| `--special` | Not required | Standard chat template for Nano variants |
### Hardware & Quantization
#### Nano-4B
| Precision | Memory |
|---|---|
| Q8_0 (8-bit, recommended) | ~3 GB |
| 4-bit | ~5 GB |
#### Nano-30B-A3B
| Precision | Memory |
|---|---|
| **UD-Q4_K_XL (4-bit, recommended)** | **~24 GB** |
| 8-bit | ~36 GB |
### Example Commands
```bash
# Nano-4B (8-bit)
./llama.cpp/llama-server \
-hf unsloth/Nemotron-3-Nano-4B-GGUF:Q8_0 \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --flash-attn on --port 8001
# Nano-30B-A3B (4-bit)
./llama.cpp/llama-server \
-hf unsloth/Nemotron-3-Nano-30B-A3B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 1.0 --min-p 0.01 \
--seed 3407 --flash-attn on --port 8001
```
---
## 3. Qwen 3.5 Family (0.8B – 397B-A17B)
### Model Variants
| Variant | Architecture | Context | Languages | Best Fit |
|---|---|---|---|---|
| **0.8B** | Dense | 256K | 201 | Smallest edge inference |
| **2B** | Dense | 256K | 201 | Small device inference |
| **4B** | Dense | 256K | 201 | Lightweight local use |
| **9B** | Dense | 256K | 201 | Capable small model |
| **27B** | Dense | 256K | 201 | Slightly more accurate than 35B-A3B; fits 18GB |
| **35B-A3B** | MoE (3B active) | 256K | 201 | Best speed/quality tradeoff; fits 22GB |
| **122B-A10B** | MoE (10B active) | 256K | 201 | High quality; needs ~70GB (4-bit) |
| **397B-A17B** | MoE (17B active) | 256K (extendable to 1M via YaRN) | 201 | Top-tier performance |
### 27B vs 35B-A3B Decision
- **27B** — Choose for slightly more accurate results when you can't fit a larger model.
- **35B-A3B** — Choose for much faster inference. MoE with only 3B active parameters.
### Hardware Requirements
| Variant | 3-bit | 4-bit | 6-bit | 8-bit | BF16 |
|---|---|---|---|---|---|
| **0.8B / 2B** | 3 GB | 3.5 GB | 5 GB | 7.5 GB | 9 GB |
| **4B** | 4.5 GB | 5.5 GB | 7 GB | 10 GB | 14 GB |
| **9B** | 5.5 GB | 6.5 GB | 9 GB | 13 GB | 19 GB |
| **27B** | 14 GB | 17 GB | 24 GB | 30 GB | 54 GB |
| **35B-A3B** | 17 GB | 22 GB | 30 GB | 38 GB | 70 GB |
| **122B-A10B** | 60 GB | 70 GB | 106 GB | 132 GB | 245 GB |
| **397B-A17B** | 180 GB | 214 GB | 340 GB | 512 GB | 810 GB |
### Inference Parameters
#### Thinking Mode
| Parameter | General Tasks | Precise Coding (e.g. WebDev) |
|---|---|---|
| `temperature` | **1.0** | **0.6** |
| `top_p` | **0.95** | **0.95** |
| `top_k` | **20** | **20** |
| `min_p` | **0.0** | **0.0** |
| `presence_penalty` | **1.5** | **0.0** |
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
#### Non-Thinking (Instruct) Mode
| Parameter | General Tasks | Reasoning Tasks |
|---|---|---|
| `temperature` | **0.7** | **1.0** |
| `top_p` | **0.8** | **0.95** |
| `top_k` | **20** | **20** |
| `min_p` | **0.0** | **0.0** |
| `presence_penalty` | **1.5** | **1.5** |
| `repetition_penalty` | **1.0** (disabled) | **1.0** (disabled) |
### Thinking Mode Control
Enable thinking:
```bash
--chat-template-kwargs '{"enable_thinking":true}'
```
Disable thinking:
```bash
--chat-template-kwargs '{"enable_thinking":false}'
```
#### Default Thinking Behavior by Variant
| Variant | Thinking Default |
|---|---|
| **0.8B, 2B, 4B, 9B** (Small) | **Disabled** — must explicitly enable |
| **27B, 35B-A3B, 122B-A10B, 397B-A17B** | **Enabled** — must explicitly disable if unwanted |
### Context Settings
| Setting | Value |
|---|---|
| Max context window | **262,144** (256K) |
| Context extension | Up to **1M** via YaRN |
| Recommended starting context | **16,384** (16K) for responsiveness |
| Adequate output length | **32,768** tokens |
### Quantization Notes
- All GGUFs use **Unsloth Dynamic 2.0** quantization — important layers upcasted to 8 or 16-bit even in 4-bit.
- Recommended starting point: **Dynamic 4-bit** (`UD-Q4_K_XL`).
- Minimum recommended: **Dynamic 2-bit** (`UD-Q2_K_XL`).
### Example Commands
```bash
# 35B-A3B — Thinking Mode (General)
./llama.cpp/llama-server \
-hf unsloth/Qwen3.5-35B-A3B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
# 9B — Thinking Enabled (small models default to disabled)
./llama.cpp/llama-server \
-hf unsloth/Qwen3.5-9B-GGUF:UD-Q4_K_XL \
--ctx-size 16384 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.00 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
```
---
## 4. Qwen3-Coder-Next (80B MoE)
### Model Overview
| Property | Value |
|---|---|
| Architecture | MoE — 80B total, **3B active parameters** |
| Max Context | **262,144** (256K) |
| Recommended Context | **32,768** for less memory use |
| Thinking Mode | **Non-reasoning only** — no `` blocks |
| Best For | Fast agentic coding, long-horizon reasoning, complex tool use |
| Performance Tier | Comparable to models with 10–20× more active parameters |
### Inference Parameters
| Parameter | Value | Notes |
|---|---|---|
| `temperature` | **1.0** | |
| `top_p` | **0.95** | |
| `top_k` | **40** | Note: higher than Qwen3.5 general |
| `min_p` | **0.01** | llama.cpp default is 0.05 — override to 0.01 |
| `repetition_penalty` | **1.0** (disabled) | Only increase if you see looping |
### Quick Reference One-Liner
```
temperature=1.0, top_p=0.95, top_k=40, min_p=0.01, repetition_penalty=1.0
```
### Hardware Requirements
| Precision | Memory Required |
|---|---|
| 3-bit (UD-IQ3_XXS) | ~34 GB |
| **4-bit (UD-Q4_K_XL)** | **~46 GB** |
| 8-bit | ~85 GB |
| BF16 | ~160 GB |
### Key Differences from Qwen3.5 General Models
- **No thinking mode** — this is a non-reasoning model; ultra-quick code responses
- **Higher `top_k`** (40 vs. 20) — broader sampling for creative code generation
- `enable_thinking` flag is not applicable
### Example Commands
```bash
# llama-server deployment
./llama.cpp/llama-server \
--model Qwen3-Coder-Next-UD-Q4_K_XL.gguf \
--alias "unsloth/Qwen3-Coder-Next" \
--seed 3407 \
--temp 1.0 --top-p 0.95 --min-p 0.01 --top-k 40 \
--ctx-size 32768 \
--flash-attn on --port 8001
```
### vLLM FP8 Dynamic (GPU Premium)
```bash
CUDA_VISIBLE_DEVICES='0,1,2,3' vllm serve unsloth/Qwen3-Coder-Next-FP8-Dynamic \
--served-model-name unsloth/Qwen3-Coder-Next \
--tensor-parallel-size 4 \
--tool-call-parser qwen3_coder \
--enable-auto-tool-choice \
--dtype bfloat16 --seed 3407 \
--max-model-len 200000 \
--gpu-memory-utilization 0.93 \
--port 8001
```
---
## 5. Google Gemma 4 Family (E2B – 31B)
### Model Variants
| Variant | Architecture | Context | Modalities | Best Fit |
|---|---|---|---|---|
| **E2B** | Dense + PLE | 128K | Text, Image, Audio | Phone / edge, ASR, speech translation |
| **E4B** | Dense + PLE | 128K | Text, Image, Audio | Laptops, fast local multimodal |
| **26B-A4B** | MoE (4B active) | 256K | Text, Image | Best speed/quality tradeoff |
| **31B** | Dense | 256K | Text, Image | Strongest performance |
### Inference Parameters (All Variants)
These are Google's default Gemma 4 parameters:
| Parameter | Value |
|---|---|
| `temperature` | **1.0** |
| `top_p` | **0.95** |
| `top_k` | **64** |
| `repetition_penalty` | **1.0** (disabled — only increase if looping) |
| End-of-sentence token | `` |
### Context Length
- **E2B / E4B**: max **128K**
- **26B-A4B / 31B**: max **256K**
- **Practical tip**: Start with **32K** for responsiveness, then increase as needed.
### Hardware Requirements
| Variant | 4-bit | 8-bit | BF16/FP16 |
|---|---|---|---|
| **E2B** | 4 GB | 5–8 GB | 10 GB |
| **E4B** | 5.5–6 GB | 9–12 GB | 16 GB |
| **26B-A4B** | 16–18 GB | 28–30 GB | 52 GB |
| **31B** | 17–20 GB | 34–38 GB | 62 GB |
### Quantization Recommendations
- **E2B / E4B** (small): prefer **Q8_0** (8-bit) for quality.
- **26B-A4B / 31B** (large): prefer **UD-Q4_K_XL** (Dynamic 4-bit) as starting point.
### 26B-A4B vs 31B Decision
- **26B-A4B** — Choose when RAM is limited. MoE with 4B active params = faster. Slight quality tradeoff.
- **31B** — Choose when you have ≥20 GB (4-bit) and want maximum quality. Slower inference.
### Thinking Mode
**Enable thinking** — add to system prompt:
```
<|think|>
You are a careful coding assistant. Explain your answer clearly.
```
Model outputs:
```
<|channel>thought
[internal reasoning]
[final answer]
```
**Disable thinking** via llama-server:
```bash
--chat-template-kwargs '{"enable_thinking":false}'
```
> **Multi-turn rule:** Only keep the final visible answer in chat history. Do NOT feed prior thought blocks back.
### Multimodal Settings
- **Images/Audio** go **before text** in prompts.
- **Video**: pass frames first, then instruction.
- **Audio** is only on E2B and E4B. Max audio: 30s. Max video: 60s (1 fps).
#### Visual Token Budgets
| Budget | Use Case |
|---|---|
| 70 / 140 | Classification, captioning, fast video |
| 280 / 560 | General multimodal chat, charts, UI |
| 1120 | OCR, document parsing, handwriting |
### Example Commands
```bash
# 26B-A4B (Dynamic 4-bit)
./llama.cpp/llama-server \
-hf unsloth/gemma-4-26B-A4B-it-GGUF:UD-Q4_K_XL \
--temp 1.0 --top-p 0.95 --top-k 64 \
--ctx-size 32768 --flash-attn on --port 8001
# 31B (Dynamic 4-bit) with thinking
./llama.cpp/llama-server \
-hf unsloth/gemma-4-31B-it-GGUF:UD-Q4_K_XL \
--temp 1.0 --top-p 0.95 --top-k 64 \
--ctx-size 32768 \
--chat-template-kwargs '{"enable_thinking":true}' \
--flash-attn on --port 8001
```
---
## 6. Quick Reference Cards
### Temperature by Model & Use Case
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|---|---|---|---|---|
| **Nemotron 3 Super** | 1.0 | — | 0.6 | — |
| **Nemotron 3 Nano** | 1.0 | — | 0.6 | — |
| **Qwen 3.5** (thinking) | 1.0 | 0.6 | — | — |
| **Qwen 3.5** (non-thinking) | — | — | — | 0.7 (general) / 1.0 (reasoning) |
| **Qwen3-Coder-Next** | 1.0 | 1.0 | — | — |
| **Gemma 4** | 1.0 | 1.0 | — | — |
### Top-P by Model & Use Case
| Model Family | General / Chat | Coding / Precise | Tool Calling | Non-Thinking |
|---|---|---|---|---|
| **Nemotron 3 Super** | 1.0 | — | 0.95 | — |
| **Nemotron 3 Nano** | 1.0 | — | 0.95 | — |
| **Qwen 3.5** (thinking) | 0.95 | 0.95 | — | — |
| **Qwen 3.5** (non-thinking) | — | — | — | 0.8 (general) / 0.95 (reasoning) |
| **Qwen3-Coder-Next** | 0.95 | 0.95 | — | — |
| **Gemma 4** | 0.95 | 0.95 | — | — |
### Top-K by Model
| Model Family | Top-K |
|---|---|
| **Nemotron 3** | Default (not specified) |
| **Qwen 3.5** | **20** |
| **Qwen3-Coder-Next** | **40** |
| **Gemma 4** | **64** |
### Min-P by Model
| Model Family | Min-P |
|---|---|
| **Nemotron 3 (all)** | **0.01** |
| **Qwen 3.5** | **0.0** |
| **Qwen3-Coder-Next** | **0.01** |
| **Gemma 4** | Default (not specified) |
### Presence Penalty by Model
| Model Family | General | Coding | Reasoning |
|---|---|---|---|
| **Nemotron 3** | Default | Default | Default |
| **Qwen 3.5** (thinking) | **1.5** | **0.0** | **1.5** |
| **Qwen 3.5** (non-thinking) | **1.5** | — | **1.5** |
| **Qwen3-Coder-Next** | Default | Default | Default |
| **Gemma 4** | Default | Default | Default |
### Repetition Penalty by Model
| Model Family | Value | Notes |
|---|---|---|
| **Nemotron 3** | Default | — |
| **Qwen 3.5** | **1.0** (disabled) | — |
| **Qwen3-Coder-Next** | **1.0** (disabled) | Only increase if looping |
| **Gemma 4** | **1.0** (disabled) | Only increase if looping |
### Seed Values
| Model Family | Recommended Seed |
|---|---|
| **Nemotron 3 (all)** | **3407** |
| **Qwen 3.5** | 3407 (optional) |
| **Qwen3-Coder-Next** | **3407** |
| **Gemma 4** | Default |
---
## 7. Critical Warnings by Model
### Nemotron 3 Super
- ⚠️ Do NOT attempt 1M context on first run — increase gradually from 16K/32K.
- ⚠️ Setting context to 1M may trigger **CUDA OOM and crash**.
- ⚠️ Router-layer fine-tuning **disabled by default** in Unsloth for MoE models.
### Qwen 3.5
- ⚠️ **No Qwen3.5 GGUF works in Ollama** due to separate mmproj vision files. Use llama.cpp-compatible backends.
- ⚠️ `presence_penalty` above 0.0 may cause **slight performance decrease**.
- ⚠️ If getting gibberish, check context length or add `--cache-type-k bf16 --cache-type-v bf16`.
- ⚠️ Redownload older GGUFs — all updated with improved quantization and tool-calling template fixes.
### Qwen3-Coder-Next
- ⚠️ **Update llama.cpp** — a previous bug in `vectorized key_gdiff` caused looping/output issues.
- ⚠️ **No Ollama support** — use llama.cpp-compatible backends.
- ⚠️ If context length too low, may see `exceeds the available context size` errors.
- ⚠️ Tool-calling improved after llama.cpp parsing fixes (Feb 19 update) — use recent version.
### Gemma 4
- ⚠️ **Do NOT use CUDA 13.2 runtime** for any GGUF — causes poor outputs.
- ⚠️ Use `llama-server` (not `llama-cli`) for thinking control — more reliable.
- ⚠️ Multi-turn: **only keep the final visible answer** in chat history. Do NOT feed prior thought blocks back.
---
## 8. Models Without KB Settings
The following model families are deployed in the Infrastructure-PFI environment but **do not have KB-documented inference parameters**. Settings for these models use general best practices or vendor defaults:
| Model | Notes |
|---|---|
| **DeepSeek R1 0528** | No KB doc — use general MoE defaults |
| **Mistral Small 3.1** | No KB doc — use vendor defaults |
| **GLM-4.7-Flash** | No KB doc — use vendor defaults |
| **GLM Steam 106B-A12B** | No KB doc — use general MoE defaults |
| **Granite 4.0 Micro** | No KB doc — use IBM defaults |
| **Kimi K2** | No KB doc — use general MoE defaults |
| **Skyfall R1 31B v4a** | No KB doc — use general defaults |
| **Hermes 4 14B** | No KB doc (download pending) — use vendor defaults |
---
## 9. PFI LiteLLM Gateway — Deployed Sampling Defaults
> **Live as of 2026-06-27** on the PFI gateway (`ana-docker:4000`; canonical config
> `eshpfi-management/stacks/litellm/conf/config.yaml`). Unlike §§1–8 (general vendor
> reference), this section is the **deployed reality** — keep it in sync when gateway
> sampling changes.
These are **overrideable defaults**: any caller that passes its own sampling param
wins; callers that omit one inherit the value below. (Verified — vLLM rejected an
out-of-range `presence_penalty=5.0`, proving per-request values reach the backend and
override the config default.) Values set per the `dvalin-smithy-dev` research pass
(provenance-cited in-thread, corroborated by §3 above). vLLM-only params (`top_k`,
`repetition_penalty`) ride in `extra_body` so LiteLLM's `drop_params` can't strip them.
| Gateway model(s) | temp | top_p | top_k | presence_penalty | repetition_penalty | Source |
|---|---|---|---|---|---|---|
| `granite-4.1-8b`, `summarizer`, `classifier` | **0** | — | — | — | — | IBM-canonical (temp 0 for inferencing) |
| `gen`, `summarizer-large`, `qwen-large`, `qwen3.5-122-a10b` (non-thinking) | **0.7** | 0.8 | 20 | **1.0** | — | Qwen3 non-thinking + operator anti-repetition |
| `gen-reasoning`, `qwen-large-reasoning`, `qwen3.5-122-a10b-reasoning` (thinking) | **0.6** | 0.95 | 20 | **1.0** | — | Qwen3 thinking |
| `qwen-image-bench`, `image-judge` | **0** | 1.0 | 1 | — | 1.05 | Qwen-Image-Bench judge reproducibility table |
| ~~`selene-1-mini-8b`~~ | — | — | — | — | — | **RETIRED 2026-08-23**; name 404s by design, not aliased |
| `chat-judge` | **0** | 1.0 | 1 | — | 1.05 | Repointed to `gen` 2026-08-23; deterministic judge profile copied from `image-judge`. The benchmark that selected `gen` ran at temperature 0 — match it. |
| `glm-5.1`, `glm-5.2`, `glm-5-turbo`, `glm-4.7`, `gen-frontier` | **1.0** | 0.95 | — | — | — | z.ai API defaults (5.x / 4.7 series) |
| `glm-4.5-air` | **0.6** | 0.95 | — | — | — | z.ai API default (4.5 series) |
| `qwen3-embedding`, `qwen3-reranker`, `reranker` | — | — | — | — | — | no sampling (embedding / rerank) |
**Notes:**
- **qwen "gen" family `presence_penalty: 1.0`** — operator-set anti-repetition for the
abliterated/NVFP4 Qwopus 122B-A10B. Qwen documents `presence_penalty` (0–2) as *the*
repetition lever; 1.0 is conservative (the §3 vendor general value is 1.5 — step up to
1.5 if loops persist). Do **not** use `repetition_penalty` for the Qwen3 family.
- **GLM (z.ai cloud) — only `temperature` + `top_p` are set.** z.ai's chat API schema
accepts no `top_k` / `min_p` / penalties, so they're deliberately not sent (would be
silently dropped). These temps match z.ai's own API defaults (explicit-over-implicit /
future-proofing).
- **Both `temp 0` values (granite, image-judge) are research-confirmed, not heuristic.**
Greedy is correct for constrained summ/classify (IBM) and for judge reproducibility
(Qwen judge card + LLM-as-judge practice). `temp 0.1` was explicitly evaluated and
rejected: it adds sampling noise without fixing loops, and *reduces* run-to-run score
consistency on the judge. If granite ever loops in production, fix via
`repetition_penalty` / `presence_penalty` / `max_tokens`, not a temperature floor.
- **`qwen-image-bench` / `image-judge` is arbo's hero-judge** (comfy-dev consumer) —
sampling changes there are a coordination item, not a unilateral gateway edit.
---
## KB Source Documents
| Document | Path in KB |
|---|---|
| Nemotron 3 Super Running Parameters | `reference/nemotron-3-super-running-parameters.md` |
| Nemotron 3 Nano Running Parameters | `reference/nemotron-3-nano-running-parameters.md` |
| Qwen 3.5 Running Parameters | `reference/qwen3.5-running-parameters.md` |
| Qwen3-Coder-Next Running Parameters | `reference/qwen3-coder-next-running-parameters.md` |
| Gemma 4 Running Parameters | `reference/gemma-4-running-parameters.md` |
---
*Document generated by Linus (Systems Architect) from AIPA Knowledge Base content curated by Atlas.*