fix(litellm): route kimi-k3 to the Kimi Code (coding) endpoint
The Heid panel plan uses Kimi's coding endpoint, not the general Moonshot API. kimi-k3 now → openai/k3 @ https://api.kimi.com/coding/v1 (KIMI_CODE_API_KEY, Vivace); the original general-endpoint entry is kept as kimi-k3-gen-api (api.moonshot.ai, MOONSHOT_API_KEY). Both verified live through the gateway. Same k3 constraints on both: temperature MUST be 1 (else 400), reasoning model (reasoning_content vs content, needs adequate max_tokens).
This commit is contained in:
@@ -35,7 +35,12 @@ VLLM_API_KEY=
|
|||||||
# (glm-5.1 / glm-5-turbo / glm-4.7 / glm-4.5-air).
|
# (glm-5.1 / glm-5-turbo / glm-4.7 / glm-4.5-air).
|
||||||
Z_AI_API_KEY=
|
Z_AI_API_KEY=
|
||||||
|
|
||||||
# Moonshot AI (Kimi) key — fronts kimi-k3 (https://api.moonshot.ai/v1). PAID.
|
# Kimi (Moonshot) keys — PAID.
|
||||||
|
# KIMI_CODE_API_KEY — the CODING endpoint (https://api.kimi.com/coding/v1),
|
||||||
|
# Vivace membership; fronts the primary `kimi-k3` arm (upstream model `k3`).
|
||||||
|
# MOONSHOT_API_KEY — the general endpoint (https://api.moonshot.ai/v1);
|
||||||
|
# fronts the `kimi-k3-gen-api` variant.
|
||||||
|
KIMI_CODE_API_KEY=
|
||||||
MOONSHOT_API_KEY=
|
MOONSHOT_API_KEY=
|
||||||
|
|
||||||
# --- Langfuse-ready (leave blank for the lean first cut) ---
|
# --- Langfuse-ready (leave blank for the lean first cut) ---
|
||||||
|
|||||||
@@ -50,6 +50,10 @@ services:
|
|||||||
# Moonshot Kimi, etc.). Paid — only gateway-keyed callers reach them,
|
# Moonshot Kimi, etc.). Paid — only gateway-keyed callers reach them,
|
||||||
# but they spend.
|
# but they spend.
|
||||||
- Z_AI_API_KEY=${Z_AI_API_KEY:-}
|
- Z_AI_API_KEY=${Z_AI_API_KEY:-}
|
||||||
|
# Kimi: KIMI_CODE_API_KEY = the coding endpoint (api.kimi.com/coding, the
|
||||||
|
# primary kimi-k3 arm); MOONSHOT_API_KEY = the general api.moonshot.ai
|
||||||
|
# endpoint (the kimi-k3-gen-api variant).
|
||||||
|
- KIMI_CODE_API_KEY=${KIMI_CODE_API_KEY:-}
|
||||||
- MOONSHOT_API_KEY=${MOONSHOT_API_KEY:-}
|
- MOONSHOT_API_KEY=${MOONSHOT_API_KEY:-}
|
||||||
# Langfuse-ready: blank until you bolt Langfuse on. Filling these +
|
# Langfuse-ready: blank until you bolt Langfuse on. Filling these +
|
||||||
# uncommenting the callback in config.yaml is the entire upgrade.
|
# uncommenting the callback in config.yaml is the entire upgrade.
|
||||||
|
|||||||
@@ -362,22 +362,34 @@ model_list:
|
|||||||
temperature: 0.6
|
temperature: 0.6
|
||||||
top_p: 0.95
|
top_p: 0.95
|
||||||
|
|
||||||
# --- Kimi K3 (Moonshot AI, cloud API) — paid passthrough fronted for unified
|
# --- Kimi K3 — CODING endpoint (Kimi Code / Vivace membership). THE PRIMARY
|
||||||
# logging alongside the local + z.ai inference. OpenAI-compatible endpoint
|
# Kimi arm the Heid cross-frontier panel plan uses. OpenAI-compatible base
|
||||||
# (https://api.moonshot.ai/v1) → openai/ provider. Flagship long-horizon
|
# https://api.kimi.com/coding/v1 → openai/ provider, upstream model id `k3`
|
||||||
# coding + knowledge model, 1M-token context (platform.kimi.ai docs). Model
|
# (1M-context; the coding lineup also carries k3-256k, kimi-for-coding,
|
||||||
# id `kimi-k3` confirmed live via /v1/models 2026-07-25 (siblings kimi-k2.6,
|
# kimi-for-coding-highspeed — ids confirmed live via /models 2026-07-25).
|
||||||
# kimi-k2.7-code, kimi-k2.7-code-highspeed — add explicit entries if wanted).
|
# PAID (Vivace subscription); key KIMI_CODE_API_KEY in .env. CONSTRAINT
|
||||||
# PAID — spends Moonshot credits; reachable by any gateway key scoped to it
|
# (verified live 2026-07-25): k3 accepts ONLY temperature=1 — any other value
|
||||||
# (the shared all-agents key spans all proxy models). Key in .env
|
# 400s ("only 1 is allowed for this model") — so it is pinned here; callers
|
||||||
# (MOONSHOT_API_KEY). CONSTRAINT (Moonshot, verified live 2026-07-25): K3
|
# must NOT override it. k3 is also a REASONING model (thinking-effort tiers
|
||||||
# ONLY accepts temperature=1 — any other value 400s ("only 1 is allowed for
|
# low/high/max per Kimi Code docs): CoT returns in `reasoning_content`, the
|
||||||
# this model"). So temperature is pinned to 1 here as the default; callers
|
# answer in `content` — give it adequate max_tokens or content returns EMPTY
|
||||||
# must NOT override it with another value. K3 is also a REASONING model:
|
# (reasoning eats a tiny budget). ---
|
||||||
# the CoT comes back in `reasoning_content`, the answer in `content` — give
|
|
||||||
# it adequate max_tokens or content returns EMPTY (reasoning eats a tiny
|
|
||||||
# budget). Verified live through the gateway 2026-07-25 (17+25→"42"). ---
|
|
||||||
- model_name: kimi-k3
|
- model_name: kimi-k3
|
||||||
|
litellm_params:
|
||||||
|
model: openai/k3
|
||||||
|
api_base: https://api.kimi.com/coding/v1
|
||||||
|
api_key: os.environ/KIMI_CODE_API_KEY
|
||||||
|
temperature: 1
|
||||||
|
model_info:
|
||||||
|
mode: chat
|
||||||
|
|
||||||
|
# --- Kimi K3 — GENERAL Moonshot API endpoint (https://api.moonshot.ai/v1),
|
||||||
|
# kept as the `-gen-api` variant. The plan uses the CODING endpoint above;
|
||||||
|
# this is the general-platform route (originally wired then demoted when the
|
||||||
|
# coding endpoint became canonical). OpenAI-compatible, upstream `kimi-k3`,
|
||||||
|
# key MOONSHOT_API_KEY. Same temperature=1 + reasoning-model constraints as
|
||||||
|
# the coding k3 (verified live through the gateway 2026-07-25, 17+25→"42"). ---
|
||||||
|
- model_name: kimi-k3-gen-api
|
||||||
litellm_params:
|
litellm_params:
|
||||||
model: openai/kimi-k3
|
model: openai/kimi-k3
|
||||||
api_base: https://api.moonshot.ai/v1
|
api_base: https://api.moonshot.ai/v1
|
||||||
|
|||||||
Reference in New Issue
Block a user