Kimi logo

MoonshotAI: Kimi K3 API

kimi-k3
Playground

Kimi K3 API is a MoonshotAI model accepting text and images across roughly 1.05M tokens. Its max-mode logic result is reported below; validate latency and token use with your own prompts before scaling. [1]

Kimi K3 API — LLM2014 logic 2026-08: Kimi-K3 (max) ranks #3 with a 67.66 median, 75.77 best score, and 1144s average time. Configurations are separate. [1][2]

MODALITIES
Price$2.7 / $13.5 / 1MCONTEXT1.0486M

Playground

Try Kimi K3 in the API Need Playground.

Provider

Current API Need routing and live API pricing. [1]

API Need10% off
Uptime
Total Context
1.0486M
Max Output
N/A
Input
$2.7/ 1M
Output
$13.5/ 1M
Cache Read
$0.27/ 1M
Cache Write
Route priced

Pricing

Pricing uses the current live API Need configuration for this model.

Official
Provider baseline
Input $3/ 1MOutput $15/ 1M
Baseline

Availability

Recent public performance data for this model.

API statusLive metrics
API Need gateway99.80% uptime

Benchmarks

Editorial benchmark data is shown when this model has a verified match. [2]Intelligence uses the source’s median-score ranking. Efficiency views are APINEED calculations, not official LLM2014 rankings. Test costs and reference prices are converted at ¥7 per US dollar; they are not APINEED selling prices. Missing values are N/A. These results describe the source’s tested configurations, not guaranteed APINEED endpoint performance.

1stGPT-5.5 (xhigh)80.2373.897.9%534s33,219$27.51$29.57
2ndGPT-5.6 Sol (xhigh)78.3769.4211.42%381s17,769$14.71$29.57
3rdKimi-K3 (max)75.7767.6610.7%1144s41,432$17.40$15.00
4thClaude Opus 5 (xhigh)71.2464.749.12%437s26,705$18.43$24.64
5thGLM-5.3 (max)74.4863.9114.19%1009s57,408$6.43$4.00
6thGLM-5.3-Flash (max)69.5560.5212.98%863s38,787$0.22$0.20
7thQwen3.8-Max (xhigh)64.6860.057.16%1297s65,135$9.38$5.14
8thDeepSeek V4 Pro 0813(max)70.1359.6314.97%1601s73,350$7.92$3.86
9thDeepSeek-V4-Flash-Vision-Exp (max)62.8658.107.57%974s79,420$2.86$1.29
10thGemini 3.7 Flash (high)66.6257.5613.6%131s26,596$2.75$3.70
11thGrok 4.6 (high)67.9356.4116.96%813s33,807$5.60$5.91
12thGemini 3.1 Pro (high)69.3555.8419.48%235s28,338$9.39$11.83
13thDeepSeek V4 Flash 0731 (max)66.3455.2316.75%881s74,657$2.69$1.29
14thGLM-5.2 (max)66.5354.7017.78%891s46,273$5.18$4.00
15thQwen3.8-Flash (xhigh)65.4454.4916.73%844s64,942$0.70$0.39
16thGPT-5.6 Luna (xhigh)65.9251.6221.69%247s37,885$1.25$1.18
17thQwen3.8-27B (xhigh)58.4847.6518.52%2318s73,987$3.55$1.71
18thDoubao-Seed-2.1-pro (high)55.1245.7816.94%2026s85,238$10.23$4.29
19thMuse Spark 1.2 (xhigh)57.9345.1422.08%359s47,209$5.54$4.19
20thGemini 3.7 Flash (low)53.5044.0617.64%59s10,673$1.10$3.70
21stClaude Sonnet 5 (xhigh)47.8039.5717.22%545s45,066$12.44$9.86
22ndTencent Hy3 (high)57.8939.5131.75%1076s53,214$0.85$0.57
23rdQwen3.7-Plus (high)51.3638.7424.57%1254s58,610$1.88$1.14
24thMiniMax-M343.8337.3714.74%982s57,486$1.93$1.20
25thDoubao-Seed-2.0-lite 0428 (high)34.1326.1823.29%669s27,845$0.40$0.51
26thClaude Opus 534.9424.5829.65%156s10,261$7.08$24.64
27thGemini 3.5 Flash Lite (high)29.5120.4430.74%192s17,842$1.23$2.46
28thLing-3.0-flash28.3717.0539.9%735s95,538N/AN/A
29thGemma 4 31B21.3616.0924.67%756s15,290$0.17$0.39
30thopenPangu-2.0-Pro20.0215.7521.33%645s28,215$1.64$2.07
31stQwen3.7-Plus20.8415.6325%323s8,910$0.29$1.14
32ndGPT-5.5 Instant21.7314.0935.16%28s1,712$1.42$29.57
33rdMistral Medium 3.517.2513.8719.59%200s28,129$5.82$7.39
34thStep-3.7-Flash26.1813.7447.52%282s47,180$1.53$1.16
35thMiMo-V2.5-Pro26.3113.0050.59%550s33,095$0.79$0.86
36thDots3-Note Preview18.6312.9530.49%210s28,386N/AN/A
37thERNIE 5.116.3211.3230.64%468s24,522$1.77$2.57
38thQwen3.8-27B16.4810.6135.62%192s8,914$0.43$1.71
39thopenPangu-2.0-Flash19.2410.1647.19%583s29,052$0.19$0.23
40thLongCat-2.017.619.7444.69%374s14,273$0.46$1.14
41stClaude Sonnet 516.499.6741.36%113s6,175$1.70$9.86
42ndGLM-5.212.428.0535.19%24s929$0.10$4.00
43rdDeepSeek V4 Flash 073113.577.9341.56%53s4,621$0.17$1.29
44thGemini 3.1 Flash Lite9.017.1121.09%11s707$0.03$1.48
45thDoubao-Seed-2.1-pro13.036.1752.65%319s6,948$0.83$4.29
46thMistral Medium 3.58.455.9329.82%18s2,589$0.54$7.39

Quick Start

Call Kimi K3 with its real backend model ID.

1

Get an API key

Create an API key in the product app.

export API_NEED_API_KEY=sk-apineed-v1-...
2

Send a request

Call /v1/chat/completions with model kimi-k3.

TypeScript SDKPythoncURLOpenAI SDK
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.API_NEED_API_KEY,
  baseURL: "https://apineed.com/v1"
});

const result = await client.chat.completions.create({
  model: "kimi-k3",
  messages: [
    { role: "user", content: "Why is the sky blue?" }
  ]
});

console.log(result.choices[0].message.content);
3

Handle the response

Validate responses and errors in your application.

curl https://apineed.com/v1/chat/completions \
  -H "Authorization: Bearer $API_NEED_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
  "model": "kimi-k3",
  "stream": true,
  "messages": [
    {
      "role": "user",
      "content": "Hello"
    }
  ]
}'

Endpoints

Endpoints currently returned by the live pricing API.

POST/v1/chat/completions
Authorization
Bearer $API_NEED_API_KEY
Content-Type
application/json
HTTP-Referer
optional - your site URL, for rankings
X-Title
optional - your site name, for rankings
Model
kimi-k3
POST/v1/responses
Authorization
Bearer $API_NEED_API_KEY
Content-Type
application/json
HTTP-Referer
optional - your site URL, for rankings
X-Title
optional - your site name, for rankings
Model
kimi-k3

Parameters

Supported parameters are shown when verified metadata is available. [1]

No verified parameter metadata is currently available.

Kimi K3 API Q&A

Model-specific answers for teams comparing Kimi K3 API on capability, benchmark evidence, integration, and cost.

What is Kimi K3 API’s verified APINEED discount?

It is 10%: $2.70 input, $13.50 output, and $0.27 cache read per million tokens where applicable.

How strong is Kimi K3 API on LLM2014 logic?

Kimi K3 — LLM2014 logic 2026-08: Kimi-K3 (max) ranks #3 with a 67.66 median, 75.77 best score, and 1144s average time. Configurations are separate.

Is Kimi K3 API suited to real-time chat?

Test interactive latency with your actual context and reasoning settings. The max-mode benchmark is a particular test configuration, not a guarantee of chat response time.

Does Kimi K3 accept visual input?

For Kimi K3 API, yes. The catalog lists text and image input, text output, and a context window of roughly 1.05M tokens.

Which endpoint serves Kimi K3 API?

Call Kimi K3 API through https://apineed.com/v1/chat/completions with kimi-k3 as the model value. Existing OpenAI SDK clients usually need only the APINEED base URL and key.

How do I validate Kimi K3 API before release?

Evaluate Kimi K3 API on representative prompts, record quality, latency, and token use, then choose reasoning settings and fallbacks from those results.

How to deploy Kimi K3

Connect the live model to your application.

1

Top up

Add usage credit in the product app.

2

Create a key

Create and securely store an API key.

3

Configure the endpoint

Send requests to /v1/chat/completions.

4

Monitor responses

Track status, latency, and errors in your application.

More live models