Kimi logo

MoonshotAI: Kimi K3 API

moonshotai/kimi-k3
Playground

Kimi K3 API is MoonshotAI’s top-ranked catalog model, accepting text and images across roughly 1.05M tokens. It places #2 in the current logic benchmark, aimed at difficult reasoning where long completion time is acceptable. [1]

Kimi K3 API posts a 74.80 median and 82.91 best score, behind only GPT-5.5 xhigh. Its 1,095-second measured completion time and 39,912 generated tokens indicate a compute-heavy max configuration. [1][2]

MODALITIES
Price$2.7 / $13.5 / 1MCONTEXT1.0486M

Playground

Test Kimi K3 API with representative production prompts before using kimi-k3 in a live workflow.

Providers

OpenRouter lists Kimi K3 API with text and image input, text output, and a 1.048576M context window. APINEED routing and prepaid rates are identified separately. [1]

API Need10% off
Uptime
Total Context
1.0486M
Max Output
N/A
Input
$2.7/ 1M
Output
$13.5/ 1M
Cache Read
$0.27/ 1M
Cache Write
Route priced

Discount

The 10% APINEED discount changes official rates of $3 input, $15 output, and $0.30 cache read to $2.70, $13.50, and $0.27. Output remains the main cost driver. The table compares official Kimi K3 API pricing with the APINEED prepaid rate.

Official
Provider baseline
Input $3/ 1MOutput $15/ 1M
Baseline

Availability

APINEED continuously monitors Kimi K3 API access and keeps requests on healthy capacity.

System statusLast 90 days
APINEED GATEWAY99.80% uptime

Kimi K3 API Benchmarks

Kimi K3 ranks #2, between GPT-5.5 xhigh and GPT-5.6 Sol. It records a stronger median than GPT-5.6 Sol but a much longer measured completion time. The complete LLM2014 logic 2026-07 table remains below, with Kimi K3 API highlighted when a current row exists. [2]

1stGPT-5.5 (xhigh)83.8077.467.57%494s30,811$25.51$29.57
2ndKimi-K3 (max)82.9174.809.78%1095s39,912$16.76$15.00
3rdClaude Opus 4.8 (xhigh)82.6266.7019.27%791s41,833$28.86$24.64
4thGPT-5.6 Sol (xhigh)81.9472.9910.92%369s17,433$14.43$29.57
5thClaude Opus 5 (xhigh)78.3871.888.29%426s26,392$18.21$24.64
6thQwen3.7-Max (xhigh)74.5666.9510.21%374s54,207$7.81$5.14
7thGLM-5.2 (max)73.6858.2720.91%890s45,683$5.12$4.00
8thGemini 3.1 Pro (high)73.3658.4620.31%214s25,905$8.58$11.83
9thGPT-5.6 Luna (xhigh)69.4951.9725.21%242s38,083$6.31$5.91
10thDeepSeek V4 Pro (max)68.0049.9726.51%1468s60,558$1.45$0.86
11thGrok 4.5 (high)67.5956.7216.08%451s43,372$7.18$5.91
12thMuse Spark 1.166.1857.2313.52%542s42,416$4.98$4.19
13thGemini 3.5 Flash (high)65.9260.398.39%188s36,372$9.03$8.87
14thDoubao-Seed-2.1-pro (high)58.6949.3515.91%2003s85,059$10.21$4.29
15thQwen3.7-Plus (high)58.5044.0924.63%1237s57,152$1.83$1.14
16thGemini 3.6 Flash (high)56.8741.7426.6%206s26,329$5.45$7.39
17thTencent Hy3 (high)54.4244.6018.04%1035s52,037$0.83$0.57
18thClaude Sonnet 5 (xhigh)51.3743.1416.02%529s44,118$12.18$9.86
19thDeepSeek V4 Flash (max)50.6236.2428.41%611s49,593$0.40$0.29
20thMiniMax-M349.1840.9416.75%985s56,467$1.90$1.20
21stClaude Opus 542.0831.7324.6%154s10,180$7.02$24.64
22ndClaude Opus 4.636.8825.7930.07%65s4,387$3.03$24.64
23rdDoubao-Seed-2.0-lite 0428 (high)35.3226.7724.21%656s26,726$0.38$0.51
24thGemini 3.5 Flash (minimal)33.2422.2233.15%44s6,189$1.54$8.87
25thLing-3.0-flash32.5320.6236.61%688s86,204$0.00$0.00
26thGemini 3.5 Flash Lite (high)30.2322.8224.51%188s18,297$1.26$2.46
27thGPT-5.5 Instant28.8717.6638.83%25s1,673$1.39$29.57
28thGemma 4 31B27.9122.0421.03%770s15,050$0.17$0.39
29thMiMo-V2.5-Pro26.9113.0051.69%478s29,443$0.71$0.86
30thQwen3.7-Plus26.1916.2338.03%314s8,692$0.28$1.14
31stStep-3.7-Flash26.1813.7447.52%258s44,912$1.46$1.16
32ndQwen3.5-27B24.9617.9827.96%451s26,391$0.51$0.69
33rdGemini 3.1 Flash Lite (high)23.7814.3039.87%82s28,351$1.17$1.48
34thClaude Sonnet 523.6410.8654.06%109s5,962$1.65$9.86
35thQwen3.7-Max22.4518.5317.46%196s6,710$0.97$5.14
36thERNIE 5.120.4915.4924.4%457s24,093$1.73$2.57
37thopenPangu-2.0-Flash19.8310.7645.74%571s28,199$0.18$0.23
38thLongCat-2.019.409.7449.79%395s14,886$0.48$1.14
39thDeepSeek V4 Flash19.3913.3731.05%63s5,286$0.04$0.29
40thMistral Medium 3.518.4413.8724.78%194s27,854$5.77$7.39
41stDoubao-Seed-2.1-pro17.499.7444.31%305s6,009$0.72$4.29
42ndGLM-5.214.8010.4329.53%23s953$0.11$4.00
43rdLing-2.6-1T13.058.3935.71%229s4,770$0.31$2.29
44thiFLYTEK Spark X212.396.2149.88%458s11,285$0.09$0.29
45thGemini 3.1 Flash Lite11.7910.2812.81%9s773$0.03$1.48
46thMistral Medium 3.510.837.1334.16%17s2,546$0.53$7.39
47thLing-2.6-flash10.074.8551.84%25s4,997$0.04$0.30

Quick Start

Connect Kimi K3 API without changing the OpenAI-style request shape. Use Kimi K3 for long-horizon reasoning and text-plus-image tasks that can run asynchronously. Set output limits and monitor token use before scaling max-mode traffic.

1

Get your API key

Create an APINEED key for Kimi K3 API and keep it in an environment variable.

export API_NEED_API_KEY=sk-apineed-v1-...
2

Make your first request

Use kimi-k3 for Kimi K3 API with the APINEED API. The request shape is compatible with OpenAI chat completions, so most SDKs only need a base URL change.

TypeScript SDKPythoncURLOpenAI SDK
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.API_NEED_API_KEY,
  baseURL: "https://apineed.com/v1"
});

const result = await client.chat.completions.create({
  model: "moonshotai/kimi-k3",
  messages: [
    { role: "user", content: "Why is the sky blue?" }
  ]
});

console.log(result.choices[0].message.content);
3

Enable streaming and fallbacks

Add stream: true when Kimi K3 API should return server-sent events. APINEED keeps routing, provider health, and fallback handling behind the same endpoint.

curl https://apineed.com/v1/chat/completions \
  -H "Authorization: Bearer $API_NEED_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "moonshotai/kimi-k3",
    "stream": true,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

Kimi K3 API Endpoint

Kimi K3 API accepts chat conversations here with streaming or non-streaming text output; the supported controls are listed on this page.

POST/v1/chat/completions
Authorization
Bearer $API_NEED_API_KEY
Content-Type
application/json
HTTP-Referer
optional - your site URL, for rankings
X-Title
optional - your site name, for rankings
Model
moonshotai/kimi-k3

Parameters

Kimi K3 API parameters currently listed for moonshotai/kimi-k3 by the OpenRouter Models API. [1]

NameTypeStatusDescription
frequency_penaltynumberSupportedPenalizes repeated token frequency.
include_reasoningbooleanSupportedIncludes reasoning content in the response when available.
logit_biasobjectSupportedAdjusts the likelihood of selected tokens.
logprobsbooleanSupportedReturns token log probabilities.
max_tokensintegerSupportedLimits generated output tokens.
min_pnumberSupportedApplies minimum-probability sampling.
presence_penaltynumberSupportedPenalizes tokens already present in the output.
reasoningobjectSupportedControls reasoning behavior and token allocation.
reasoning_effortenumSupportedSelects the model reasoning effort.
repetition_penaltynumberSupportedControls repetition across generated tokens.
response_formatobjectSupportedRequests a specific response format.
seedintegerSupportedRequests deterministic sampling when supported by the provider.
stopstring or arraySupportedStops generation at the supplied sequence.
structured_outputsbooleanSupportedEnables schema-constrained structured output.
temperaturenumberSupportedControls sampling randomness.
tool_choicestring or objectSupportedControls which tool the model may call.
toolsarraySupportedDefines tools available to the model.
top_kintegerSupportedRestricts sampling to the highest-probability tokens.
top_logprobsintegerSupportedSets how many top-token log probabilities are returned.
top_pnumberSupportedControls nucleus sampling.

Kimi K3 API Q&A

Model-specific answers for teams comparing Kimi K3 API on capability, benchmark evidence, integration, and cost.

What is Kimi K3 API’s verified APINEED discount?

It is 10%: $2.70 input, $13.50 output, and $0.27 cache read per million tokens where applicable.

How strong is Kimi K3 API on LLM2014 logic?

Its max run ranks #2 with a 74.80 median and 82.91 best score in the July 2026 dataset.

Is Kimi K3 API suited to real-time chat?

Its benchmark run took 1,095 seconds on average, so teams should test interactive latency rather than assume max mode is conversational.

Does Kimi K3 accept visual input?

For Kimi K3 API, yes. The catalog lists text and image input, text output, and a context window of roughly 1.05M tokens.

Which endpoint serves Kimi K3 API?

Call Kimi K3 API through https://apineed.com/v1/chat/completions with kimi-k3 as the model value. Existing OpenAI SDK clients usually need only the APINEED base URL and key.

How do I validate Kimi K3 API before release?

Evaluate Kimi K3 API on representative prompts, record quality, latency, and token use, then choose reasoning settings and fallbacks from those results.

How to Deploy the Kimi K3 API on apineed.com

Deploy Kimi K3 API through APINEED after validating its model-specific trade-offs above. The API key, credit, and endpoint flow stays consistent across the catalog.

1

Add credits for Kimi K3 API

Add credits on APINEED before deploying the Kimi K3 API. Pay-as-you-go billing lets usage start small and scale with production traffic.

2

Get your API key

Create an APINEED API key for the Kimi K3 API. The same key can call Kimi K3 and other AI APIs through apineed.com.

3

Set the model slug

Use kimi-k3 as the model value when you deploy the Kimi K3 API. Keep the APINEED base URL at https://apineed.com/v1.

4

Send a request to Kimi K3 API

Send chat completions or responses to the Kimi K3 API from Claude Code, Codex, or any custom agent. The request keeps its OpenAI-compatible shape.

More models from MoonshotAI