Recommended Model Inference Settings — Reference Document
Source: AIPA Knowledge Base — compiled from 5 KB reference documents.
Status: Reference document (KB-derived; freshness not actively maintained — verify against current upstream model docs before relying on a setting for production work).
Purpose: Canonical reference for llama-server / llama.cpp inference parameters across all model families with KB-documented settings.
Table of Contents
- NVIDIA Nemotron 3 Super (120B-A12B)
- NVIDIA Nemotron 3 Nano (4B / 30B-A3B)
- Qwen 3.5 Family (0.8B – 397B-A17B)
- Qwen3-Coder-Next (80B MoE)
- Google Gemma 4 Family (E2B – 31B)
- Quick Reference Cards
- Critical Warnings by Model
- Models Without KB Settings
- PFI LiteLLM Gateway — Deployed Sampling Defaults
1. NVIDIA Nemotron 3 Super (120B-A12B)
Model Overview
| Property |
Value |
| Architecture |
MoE — 120B total, 12B active parameters |
| Max Context |
1,048,576 (1M tokens) |
| Recommended Starting Context |
16K or 32K — increase gradually |
| Reasoning Tokens |
<think)> (ID 12), </think)> (ID 13) |
| Positional Embeddings |
NoPE — YaRN NOT needed |
| Best For |
Multi-agent AI, high-efficiency reasoning, coding, math |
| Performance Tier |
~GPT-5.2 / Claude Opus 4.5 level |
Inference Parameters
| Parameter |
General Chat / Instruction |
Tool Calling |
temperature |
1.0 |
0.6 |
top_p |
1.0 |
0.95 |
min_p |
0.01 |
0.01 |
Additional Settings
| Setting |
Value |
Notes |
--seed |
3407 |
Reproducibility |
--prio |
2 or 3 |
Priority scheduling |
--special |
Required |
To see reasoning tokens |
--verbose-prompt |
Required |
To see prepended <think)> tokens |
max_new_tokens |
32,768 – 262,144 |
Up to 1M |
Quantization & Memory
| Precision |
Memory Required |
| UD-Q2_K_XL (2-bit) |
~32–36 GB |
| UD-Q4_K_XL (4-bit) |
~64–72 GB |
| 8-bit |
~128 GB |
| BF16 |
~240 GB |
Example Command
2. NVIDIA Nemotron 3 Nano (4B / 30B-A3B)
Model Variants
| Variant |
Architecture |
Context |
Active Params |
Best Fit |
| Nano-4B |
Dense |
128K |
4B |
Lightweight coding, math, agentic tasks |
| Nano-30B-A3B |
MoE |
128K |
3B active |
Best performance/size on 24GB devices |
Inference Parameters (Both Variants)
| Parameter |
General Chat / Instruction |
Tool Calling |
temperature |
1.0 |
0.6 |
top_p |
1.0 |
0.95 |
min_p |
0.01 |
0.01 |
Additional Settings
| Setting |
Value |
Notes |
--seed |
3407 |
Reproducibility |
--special |
Not required |
Standard chat template for Nano variants |
Hardware & Quantization
Nano-4B
| Precision |
Memory |
| Q8_0 (8-bit, recommended) |
~3 GB |
| 4-bit |
~5 GB |
Nano-30B-A3B
| Precision |
Memory |
| UD-Q4_K_XL (4-bit, recommended) |
~24 GB |
| 8-bit |
~36 GB |
Example Commands
3. Qwen 3.5 Family (0.8B – 397B-A17B)
Model Variants
| Variant |
Architecture |
Context |
Languages |
Best Fit |
| 0.8B |
Dense |
256K |
201 |
Smallest edge inference |
| 2B |
Dense |
256K |
201 |
Small device inference |
| 4B |
Dense |
256K |
201 |
Lightweight local use |
| 9B |
Dense |
256K |
201 |
Capable small model |
| 27B |
Dense |
256K |
201 |
Slightly more accurate than 35B-A3B; fits 18GB |
| 35B-A3B |
MoE (3B active) |
256K |
201 |
Best speed/quality tradeoff; fits 22GB |
| 122B-A10B |
MoE (10B active) |
256K |
201 |
High quality; needs ~70GB (4-bit) |
| 397B-A17B |
MoE (17B active) |
256K (extendable to 1M via YaRN) |
201 |
Top-tier performance |
27B vs 35B-A3B Decision
- 27B — Choose for slightly more accurate results when you can't fit a larger model.
- 35B-A3B — Choose for much faster inference. MoE with only 3B active parameters.
Hardware Requirements
| Variant |
3-bit |
4-bit |
6-bit |
8-bit |
BF16 |
| 0.8B / 2B |
3 GB |
3.5 GB |
5 GB |
7.5 GB |
9 GB |
| 4B |
4.5 GB |
5.5 GB |
7 GB |
10 GB |
14 GB |
| 9B |
5.5 GB |
6.5 GB |
9 GB |
13 GB |
19 GB |
| 27B |
14 GB |
17 GB |
24 GB |
30 GB |
54 GB |
| 35B-A3B |
17 GB |
22 GB |
30 GB |
38 GB |
70 GB |
| 122B-A10B |
60 GB |
70 GB |
106 GB |
132 GB |
245 GB |
| 397B-A17B |
180 GB |
214 GB |
340 GB |
512 GB |
810 GB |
Inference Parameters
Thinking Mode
| Parameter |
General Tasks |
Precise Coding (e.g. WebDev) |
temperature |
1.0 |
0.6 |
top_p |
0.95 |
0.95 |
top_k |
20 |
20 |
min_p |
0.0 |
0.0 |
presence_penalty |
1.5 |
0.0 |
repetition_penalty |
1.0 (disabled) |
1.0 (disabled) |
Non-Thinking (Instruct) Mode
| Parameter |
General Tasks |
Reasoning Tasks |
temperature |
0.7 |
1.0 |
top_p |
0.8 |
0.95 |
top_k |
20 |
20 |
min_p |
0.0 |
0.0 |
presence_penalty |
1.5 |
1.5 |
repetition_penalty |
1.0 (disabled) |
1.0 (disabled) |
Thinking Mode Control
Enable thinking:
Disable thinking:
Default Thinking Behavior by Variant
| Variant |
Thinking Default |
| 0.8B, 2B, 4B, 9B (Small) |
Disabled — must explicitly enable |
| 27B, 35B-A3B, 122B-A10B, 397B-A17B |
Enabled — must explicitly disable if unwanted |
Context Settings
| Setting |
Value |
| Max context window |
262,144 (256K) |
| Context extension |
Up to 1M via YaRN |
| Recommended starting context |
16,384 (16K) for responsiveness |
| Adequate output length |
32,768 tokens |
Quantization Notes
- All GGUFs use Unsloth Dynamic 2.0 quantization — important layers upcasted to 8 or 16-bit even in 4-bit.
- Recommended starting point: Dynamic 4-bit (
UD-Q4_K_XL).
- Minimum recommended: Dynamic 2-bit (
UD-Q2_K_XL).
Example Commands
4. Qwen3-Coder-Next (80B MoE)
Model Overview
| Property |
Value |
| Architecture |
MoE — 80B total, 3B active parameters |
| Max Context |
262,144 (256K) |
| Recommended Context |
32,768 for less memory use |
| Thinking Mode |
Non-reasoning only — no <think)> blocks |
| Best For |
Fast agentic coding, long-horizon reasoning, complex tool use |
| Performance Tier |
Comparable to models with 10–20× more active parameters |
Inference Parameters
| Parameter |
Value |
Notes |
temperature |
1.0 |
|
top_p |
0.95 |
|
top_k |
40 |
Note: higher than Qwen3.5 general |
min_p |
0.01 |
llama.cpp default is 0.05 — override to 0.01 |
repetition_penalty |
1.0 (disabled) |
Only increase if you see looping |
Quick Reference One-Liner
Hardware Requirements
| Precision |
Memory Required |
| 3-bit (UD-IQ3_XXS) |
~34 GB |
| 4-bit (UD-Q4_K_XL) |
~46 GB |
| 8-bit |
~85 GB |
| BF16 |
~160 GB |
Key Differences from Qwen3.5 General Models
- No thinking mode — this is a non-reasoning model; ultra-quick code responses
- Higher
top_k (40 vs. 20) — broader sampling for creative code generation
enable_thinking flag is not applicable
Example Commands
vLLM FP8 Dynamic (GPU Premium)
5. Google Gemma 4 Family (E2B – 31B)
Model Variants
| Variant |
Architecture |
Context |
Modalities |
Best Fit |
| E2B |
Dense + PLE |
128K |
Text, Image, Audio |
Phone / edge, ASR, speech translation |
| E4B |
Dense + PLE |
128K |
Text, Image, Audio |
Laptops, fast local multimodal |
| 26B-A4B |
MoE (4B active) |
256K |
Text, Image |
Best speed/quality tradeoff |
| 31B |
Dense |
256K |
Text, Image |
Strongest performance |
Inference Parameters (All Variants)
These are Google's default Gemma 4 parameters:
| Parameter |
Value |
temperature |
1.0 |
top_p |
0.95 |
top_k |
64 |
repetition_penalty |
1.0 (disabled — only increase if looping) |
| End-of-sentence token |
`<turn |
Context Length
- E2B / E4B: max 128K
- 26B-A4B / 31B: max 256K
- Practical tip: Start with 32K for responsiveness, then increase as needed.
Hardware Requirements
| Variant |
4-bit |
8-bit |
BF16/FP16 |
| E2B |
4 GB |
5–8 GB |
10 GB |
| E4B |
5.5–6 GB |
9–12 GB |
16 GB |
| 26B-A4B |
16–18 GB |
28–30 GB |
52 GB |
| 31B |
17–20 GB |
34–38 GB |
62 GB |
Quantization Recommendations
- E2B / E4B (small): prefer Q8_0 (8-bit) for quality.
- 26B-A4B / 31B (large): prefer UD-Q4_K_XL (Dynamic 4-bit) as starting point.
26B-A4B vs 31B Decision
- 26B-A4B — Choose when RAM is limited. MoE with 4B active params = faster. Slight quality tradeoff.
- 31B — Choose when you have ≥20 GB (4-bit) and want maximum quality. Slower inference.
Thinking Mode
Enable thinking — add to system prompt:
Model outputs:
Disable thinking via llama-server:
Multi-turn rule: Only keep the final visible answer in chat history. Do NOT feed prior thought blocks back.
Multimodal Settings
- Images/Audio go before text in prompts.
- Video: pass frames first, then instruction.
- Audio is only on E2B and E4B. Max audio: 30s. Max video: 60s (1 fps).
Visual Token Budgets
| Budget |
Use Case |
| 70 / 140 |
Classification, captioning, fast video |
| 280 / 560 |
General multimodal chat, charts, UI |
| 1120 |
OCR, document parsing, handwriting |
Example Commands
6. Quick Reference Cards
Temperature by Model & Use Case
| Model Family |
General / Chat |
Coding / Precise |
Tool Calling |
Non-Thinking |
| Nemotron 3 Super |
1.0 |
— |
0.6 |
— |
| Nemotron 3 Nano |
1.0 |
— |
0.6 |
— |
| Qwen 3.5 (thinking) |
1.0 |
0.6 |
— |
— |
| Qwen 3.5 (non-thinking) |
— |
— |
— |
0.7 (general) / 1.0 (reasoning) |
| Qwen3-Coder-Next |
1.0 |
1.0 |
— |
— |
| Gemma 4 |
1.0 |
1.0 |
— |
— |
Top-P by Model & Use Case
| Model Family |
General / Chat |
Coding / Precise |
Tool Calling |
Non-Thinking |
| Nemotron 3 Super |
1.0 |
— |
0.95 |
— |
| Nemotron 3 Nano |
1.0 |
— |
0.95 |
— |
| Qwen 3.5 (thinking) |
0.95 |
0.95 |
— |
— |
| Qwen 3.5 (non-thinking) |
— |
— |
— |
0.8 (general) / 0.95 (reasoning) |
| Qwen3-Coder-Next |
0.95 |
0.95 |
— |
— |
| Gemma 4 |
0.95 |
0.95 |
— |
— |
Top-K by Model
| Model Family |
Top-K |
| Nemotron 3 |
Default (not specified) |
| Qwen 3.5 |
20 |
| Qwen3-Coder-Next |
40 |
| Gemma 4 |
64 |
Min-P by Model
| Model Family |
Min-P |
| Nemotron 3 (all) |
0.01 |
| Qwen 3.5 |
0.0 |
| Qwen3-Coder-Next |
0.01 |
| Gemma 4 |
Default (not specified) |
Presence Penalty by Model
| Model Family |
General |
Coding |
Reasoning |
| Nemotron 3 |
Default |
Default |
Default |
| Qwen 3.5 (thinking) |
1.5 |
0.0 |
1.5 |
| Qwen 3.5 (non-thinking) |
1.5 |
— |
1.5 |
| Qwen3-Coder-Next |
Default |
Default |
Default |
| Gemma 4 |
Default |
Default |
Default |
Repetition Penalty by Model
| Model Family |
Value |
Notes |
| Nemotron 3 |
Default |
— |
| Qwen 3.5 |
1.0 (disabled) |
— |
| Qwen3-Coder-Next |
1.0 (disabled) |
Only increase if looping |
| Gemma 4 |
1.0 (disabled) |
Only increase if looping |
Seed Values
| Model Family |
Recommended Seed |
| Nemotron 3 (all) |
3407 |
| Qwen 3.5 |
3407 (optional) |
| Qwen3-Coder-Next |
3407 |
| Gemma 4 |
Default |
7. Critical Warnings by Model
Nemotron 3 Super
- ⚠️ Do NOT attempt 1M context on first run — increase gradually from 16K/32K.
- ⚠️ Setting context to 1M may trigger CUDA OOM and crash.
- ⚠️ Router-layer fine-tuning disabled by default in Unsloth for MoE models.
Qwen 3.5
- ⚠️ No Qwen3.5 GGUF works in Ollama due to separate mmproj vision files. Use llama.cpp-compatible backends.
- ⚠️
presence_penalty above 0.0 may cause slight performance decrease.
- ⚠️ If getting gibberish, check context length or add
--cache-type-k bf16 --cache-type-v bf16.
- ⚠️ Redownload older GGUFs — all updated with improved quantization and tool-calling template fixes.
Qwen3-Coder-Next
- ⚠️ Update llama.cpp — a previous bug in
vectorized key_gdiff caused looping/output issues.
- ⚠️ No Ollama support — use llama.cpp-compatible backends.
- ⚠️ If context length too low, may see
exceeds the available context size errors.
- ⚠️ Tool-calling improved after llama.cpp parsing fixes (Feb 19 update) — use recent version.
Gemma 4
- ⚠️ Do NOT use CUDA 13.2 runtime for any GGUF — causes poor outputs.
- ⚠️ Use
llama-server (not llama-cli) for thinking control — more reliable.
- ⚠️ Multi-turn: only keep the final visible answer in chat history. Do NOT feed prior thought blocks back.
8. Models Without KB Settings
The following model families are deployed in the Infrastructure-PFI environment but do not have KB-documented inference parameters. Settings for these models use general best practices or vendor defaults:
| Model |
Notes |
| DeepSeek R1 0528 |
No KB doc — use general MoE defaults |
| Mistral Small 3.1 |
No KB doc — use vendor defaults |
| GLM-4.7-Flash |
No KB doc — use vendor defaults |
| GLM Steam 106B-A12B |
No KB doc — use general MoE defaults |
| Granite 4.0 Micro |
No KB doc — use IBM defaults |
| Kimi K2 |
No KB doc — use general MoE defaults |
| Skyfall R1 31B v4a |
No KB doc — use general defaults |
| Hermes 4 14B |
No KB doc (download pending) — use vendor defaults |
9. PFI LiteLLM Gateway — Deployed Sampling Defaults
Live as of 2026-06-27 on the PFI gateway (ana-docker:4000; canonical config
eshpfi-management/stacks/litellm/conf/config.yaml). Unlike §§1–8 (general vendor
reference), this section is the deployed reality — keep it in sync when gateway
sampling changes.
These are overrideable defaults: any caller that passes its own sampling param
wins; callers that omit one inherit the value below. (Verified — vLLM rejected an
out-of-range presence_penalty=5.0, proving per-request values reach the backend and
override the config default.) Values set per the dvalin-smithy-dev research pass
(provenance-cited in-thread, corroborated by §3 above). vLLM-only params (top_k,
repetition_penalty) ride in extra_body so LiteLLM's drop_params can't strip them.
| Gateway model(s) |
temp |
top_p |
top_k |
presence_penalty |
repetition_penalty |
Source |
granite-4.1-8b, summarizer, classifier |
0 |
— |
— |
— |
— |
IBM-canonical (temp 0 for inferencing) |
gen, summarizer-large, qwen-large, qwen3.5-122-a10b (non-thinking) |
0.7 |
0.8 |
20 |
1.0 |
— |
Qwen3 non-thinking + operator anti-repetition |
gen-reasoning, qwen-large-reasoning, qwen3.5-122-a10b-reasoning (thinking) |
0.6 |
0.95 |
20 |
1.0 |
— |
Qwen3 thinking |
qwen-image-bench, image-judge |
0 |
1.0 |
1 |
— |
1.05 |
Qwen-Image-Bench judge reproducibility table |
selene-1-mini-8b, chat-judge |
0.6 |
0.9 |
— |
— |
— |
Selene generation_config |
glm-5.1, glm-5.2, glm-5-turbo, glm-4.7, gen-frontier |
1.0 |
0.95 |
— |
— |
— |
z.ai API defaults (5.x / 4.7 series) |
glm-4.5-air |
0.6 |
0.95 |
— |
— |
— |
z.ai API default (4.5 series) |
qwen3-embedding, qwen3-reranker, reranker |
— |
— |
— |
— |
— |
no sampling (embedding / rerank) |
Notes:
- qwen "gen" family
presence_penalty: 1.0 — operator-set anti-repetition for the
abliterated/NVFP4 Qwopus 122B-A10B. Qwen documents presence_penalty (0–2) as the
repetition lever; 1.0 is conservative (the §3 vendor general value is 1.5 — step up to
1.5 if loops persist). Do not use repetition_penalty for the Qwen3 family.
- GLM (z.ai cloud) — only
temperature + top_p are set. z.ai's chat API schema
accepts no top_k / min_p / penalties, so they're deliberately not sent (would be
silently dropped). These temps match z.ai's own API defaults (explicit-over-implicit /
future-proofing).
- Both
temp 0 values (granite, image-judge) are research-confirmed, not heuristic.
Greedy is correct for constrained summ/classify (IBM) and for judge reproducibility
(Qwen judge card + LLM-as-judge practice). temp 0.1 was explicitly evaluated and
rejected: it adds sampling noise without fixing loops, and reduces run-to-run score
consistency on the judge. If granite ever loops in production, fix via
repetition_penalty / presence_penalty / max_tokens, not a temperature floor.
qwen-image-bench / image-judge is arbo's hero-judge (comfy-dev consumer) —
sampling changes there are a coordination item, not a unilateral gateway edit.
KB Source Documents
| Document |
Path in KB |
| Nemotron 3 Super Running Parameters |
reference/nemotron-3-super-running-parameters.md |
| Nemotron 3 Nano Running Parameters |
reference/nemotron-3-nano-running-parameters.md |
| Qwen 3.5 Running Parameters |
reference/qwen3.5-running-parameters.md |
| Qwen3-Coder-Next Running Parameters |
reference/qwen3-coder-next-running-parameters.md |
| Gemma 4 Running Parameters |
reference/gemma-4-running-parameters.md |
Document generated by Linus (Systems Architect) from AIPA Knowledge Base content curated by Atlas.