fix(litellm): route kimi-k3 to the Kimi Code (coding) endpoint

The Heid panel plan uses Kimi's coding endpoint, not the general Moonshot API.
kimi-k3 now → openai/k3 @ https://api.kimi.com/coding/v1 (KIMI_CODE_API_KEY,
Vivace); the original general-endpoint entry is kept as kimi-k3-gen-api
(api.moonshot.ai, MOONSHOT_API_KEY). Both verified live through the gateway.
Same k3 constraints on both: temperature MUST be 1 (else 400), reasoning model
(reasoning_content vs content, needs adequate max_tokens).
This commit is contained in:
vh
2026-07-25 11:00:15 -07:00
parent 0b0c915dc9
commit 9e2f787567
3 changed files with 37 additions and 16 deletions
+27 -15
View File
@@ -362,22 +362,34 @@ model_list:
temperature: 0.6
top_p: 0.95
# --- Kimi K3 (Moonshot AI, cloud API) — paid passthrough fronted for unified
# logging alongside the local + z.ai inference. OpenAI-compatible endpoint
# (https://api.moonshot.ai/v1) → openai/ provider. Flagship long-horizon
# coding + knowledge model, 1M-token context (platform.kimi.ai docs). Model
# id `kimi-k3` confirmed live via /v1/models 2026-07-25 (siblings kimi-k2.6,
# kimi-k2.7-code, kimi-k2.7-code-highspeed — add explicit entries if wanted).
# PAID — spends Moonshot credits; reachable by any gateway key scoped to it
# (the shared all-agents key spans all proxy models). Key in .env
# (MOONSHOT_API_KEY). CONSTRAINT (Moonshot, verified live 2026-07-25): K3
# ONLY accepts temperature=1 — any other value 400s ("only 1 is allowed for
# this model"). So temperature is pinned to 1 here as the default; callers
# must NOT override it with another value. K3 is also a REASONING model:
# the CoT comes back in `reasoning_content`, the answer in `content` — give
# it adequate max_tokens or content returns EMPTY (reasoning eats a tiny
# budget). Verified live through the gateway 2026-07-25 (17+25→"42"). ---
# --- Kimi K3 — CODING endpoint (Kimi Code / Vivace membership). THE PRIMARY
# Kimi arm the Heid cross-frontier panel plan uses. OpenAI-compatible base
# https://api.kimi.com/coding/v1 → openai/ provider, upstream model id `k3`
# (1M-context; the coding lineup also carries k3-256k, kimi-for-coding,
# kimi-for-coding-highspeed — ids confirmed live via /models 2026-07-25).
# PAID (Vivace subscription); key KIMI_CODE_API_KEY in .env. CONSTRAINT
# (verified live 2026-07-25): k3 accepts ONLY temperature=1 — any other value
# 400s ("only 1 is allowed for this model") — so it is pinned here; callers
# must NOT override it. k3 is also a REASONING model (thinking-effort tiers
# low/high/max per Kimi Code docs): CoT returns in `reasoning_content`, the
# answer in `content` — give it adequate max_tokens or content returns EMPTY
# (reasoning eats a tiny budget). ---
- model_name: kimi-k3
litellm_params:
model: openai/k3
api_base: https://api.kimi.com/coding/v1
api_key: os.environ/KIMI_CODE_API_KEY
temperature: 1
model_info:
mode: chat
# --- Kimi K3 — GENERAL Moonshot API endpoint (https://api.moonshot.ai/v1),
# kept as the `-gen-api` variant. The plan uses the CODING endpoint above;
# this is the general-platform route (originally wired then demoted when the
# coding endpoint became canonical). OpenAI-compatible, upstream `kimi-k3`,
# key MOONSHOT_API_KEY. Same temperature=1 + reasoning-model constraints as
# the coding k3 (verified live through the gateway 2026-07-25, 17+25→"42"). ---
- model_name: kimi-k3-gen-api
litellm_params:
model: openai/kimi-k3
api_base: https://api.moonshot.ai/v1