Deepseek logo

DeepSeek: DeepSeek V4 Flash API

deepseek-v4-flash
Playground

DeepSeek V4 Flash API is the lower-cost, text-only companion to DeepSeek V4 Pro. Its 1.05M context, reasoning controls, tools, and structured output support make it suitable for high-volume agent steps. [1]

The max row ranks #20 with a 36.24 median and 611-second measured time. The standard row ranks #36 but completes in 63 seconds, exposing a clear quality-versus-latency choice. [1][2]

MODALITIES
Price$0.0675 / $0.135 / 1MCONTEXT1.0486M

Playground

Test DeepSeek V4 Flash API with representative production prompts before using deepseek-v4-flash in a live workflow.

Providers

OpenRouter lists DeepSeek V4 Flash API with text input, text output, and a 1.048576M context window. APINEED routing and prepaid rates are identified separately. [1]

API Need25% off
Uptime
Total Context
1.0486M
Max Output
393.216K
Input
$0.0675/ 1M
Output
$0.135/ 1M
Cache Read
$0.0135/ 1M
Cache Write
Route priced

Discount

Official prices of $0.09 input, $0.18 output, and $0.018 cache read fall to $0.0675, $0.135, and $0.0135 with APINEED’s 25% discount. The table compares official DeepSeek V4 Flash API pricing with the APINEED prepaid rate.

Official
Provider baseline
Input $0.09/ 1MOutput $0.18/ 1M
Baseline

Availability

APINEED continuously monitors DeepSeek V4 Flash API access and keeps requests on healthy capacity.

System statusLast 90 days
APINEED GATEWAY99.19% uptime

DeepSeek V4 Flash API Benchmarks

DeepSeek V4 Flash ranks #20 in max mode and #36 in standard mode. The standard run is much faster, while max mode raises the median by 22.87 points. The complete LLM2014 logic 2026-07 table remains below, with DeepSeek V4 Flash API highlighted when a current row exists. [2]

1stGPT-5.5 (xhigh)83.8077.467.57%494s30,811$25.51$29.57
2ndKimi-K3 (max)82.9174.809.78%1095s39,912$16.76$15.00
3rdClaude Opus 4.8 (xhigh)82.6266.7019.27%791s41,833$28.86$24.64
4thGPT-5.6 Sol (xhigh)81.9472.9910.92%369s17,433$14.43$29.57
5thClaude Opus 5 (xhigh)78.3871.888.29%426s26,392$18.21$24.64
6thQwen3.7-Max (xhigh)74.5666.9510.21%374s54,207$7.81$5.14
7thGLM-5.2 (max)73.6858.2720.91%890s45,683$5.12$4.00
8thGemini 3.1 Pro (high)73.3658.4620.31%214s25,905$8.58$11.83
9thGPT-5.6 Luna (xhigh)69.4951.9725.21%242s38,083$6.31$5.91
10thDeepSeek V4 Pro (max)68.0049.9726.51%1468s60,558$1.45$0.86
11thGrok 4.5 (high)67.5956.7216.08%451s43,372$7.18$5.91
12thMuse Spark 1.166.1857.2313.52%542s42,416$4.98$4.19
13thGemini 3.5 Flash (high)65.9260.398.39%188s36,372$9.03$8.87
14thDoubao-Seed-2.1-pro (high)58.6949.3515.91%2003s85,059$10.21$4.29
15thQwen3.7-Plus (high)58.5044.0924.63%1237s57,152$1.83$1.14
16thGemini 3.6 Flash (high)56.8741.7426.6%206s26,329$5.45$7.39
17thTencent Hy3 (high)54.4244.6018.04%1035s52,037$0.83$0.57
18thClaude Sonnet 5 (xhigh)51.3743.1416.02%529s44,118$12.18$9.86
19thDeepSeek V4 Flash (max)50.6236.2428.41%611s49,593$0.40$0.29
20thMiniMax-M349.1840.9416.75%985s56,467$1.90$1.20
21stClaude Opus 542.0831.7324.6%154s10,180$7.02$24.64
22ndClaude Opus 4.636.8825.7930.07%65s4,387$3.03$24.64
23rdDoubao-Seed-2.0-lite 0428 (high)35.3226.7724.21%656s26,726$0.38$0.51
24thGemini 3.5 Flash (minimal)33.2422.2233.15%44s6,189$1.54$8.87
25thLing-3.0-flash32.5320.6236.61%688s86,204$0.00$0.00
26thGemini 3.5 Flash Lite (high)30.2322.8224.51%188s18,297$1.26$2.46
27thGPT-5.5 Instant28.8717.6638.83%25s1,673$1.39$29.57
28thGemma 4 31B27.9122.0421.03%770s15,050$0.17$0.39
29thMiMo-V2.5-Pro26.9113.0051.69%478s29,443$0.71$0.86
30thQwen3.7-Plus26.1916.2338.03%314s8,692$0.28$1.14
31stStep-3.7-Flash26.1813.7447.52%258s44,912$1.46$1.16
32ndQwen3.5-27B24.9617.9827.96%451s26,391$0.51$0.69
33rdGemini 3.1 Flash Lite (high)23.7814.3039.87%82s28,351$1.17$1.48
34thClaude Sonnet 523.6410.8654.06%109s5,962$1.65$9.86
35thQwen3.7-Max22.4518.5317.46%196s6,710$0.97$5.14
36thERNIE 5.120.4915.4924.4%457s24,093$1.73$2.57
37thopenPangu-2.0-Flash19.8310.7645.74%571s28,199$0.18$0.23
38thLongCat-2.019.409.7449.79%395s14,886$0.48$1.14
39thDeepSeek V4 Flash19.3913.3731.05%63s5,286$0.04$0.29
40thMistral Medium 3.518.4413.8724.78%194s27,854$5.77$7.39
41stDoubao-Seed-2.1-pro17.499.7444.31%305s6,009$0.72$4.29
42ndGLM-5.214.8010.4329.53%23s953$0.11$4.00
43rdLing-2.6-1T13.058.3935.71%229s4,770$0.31$2.29
44thiFLYTEK Spark X212.396.2149.88%458s11,285$0.09$0.29
45thGemini 3.1 Flash Lite11.7910.2812.81%9s773$0.03$1.48
46thMistral Medium 3.510.837.1334.16%17s2,546$0.53$7.39
47thLing-2.6-flash10.074.8551.84%25s4,997$0.04$0.30

Quick Start

Connect DeepSeek V4 Flash API without changing the OpenAI-style request shape. Route classification, extraction, and routine coding turns to standard mode, then reserve max reasoning for harder cases. Explicit output limits help control long reasoning traces.

1

Get your API key

Create an APINEED key for DeepSeek V4 Flash API and keep it in an environment variable.

export API_NEED_API_KEY=sk-apineed-v1-...
2

Make your first request

Use deepseek-v4-flash for DeepSeek V4 Flash API with the APINEED API. The request shape is compatible with OpenAI chat completions, so most SDKs only need a base URL change.

TypeScript SDKPythoncURLOpenAI SDK
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.API_NEED_API_KEY,
  baseURL: "https://apineed.com/v1"
});

const result = await client.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "user", content: "Why is the sky blue?" }
  ]
});

console.log(result.choices[0].message.content);
3

Enable streaming and fallbacks

Add stream: true when DeepSeek V4 Flash API should return server-sent events. APINEED keeps routing, provider health, and fallback handling behind the same endpoint.

curl https://apineed.com/v1/chat/completions \
  -H "Authorization: Bearer $API_NEED_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-v4-flash",
    "stream": true,
    "messages": [{ "role": "user", "content": "Hello" }]
  }'

DeepSeek V4 Flash API Endpoint

DeepSeek V4 Flash API accepts chat conversations here with streaming or non-streaming text output; the supported controls are listed on this page.

POST/v1/chat/completions
Authorization
Bearer $API_NEED_API_KEY
Content-Type
application/json
HTTP-Referer
optional - your site URL, for rankings
X-Title
optional - your site name, for rankings
Model
deepseek-v4-flash

Parameters

DeepSeek V4 Flash API parameters currently listed for deepseek-v4-flash by the OpenRouter Models API. [1]

NameTypeStatusDescription
frequency_penaltynumberSupportedPenalizes repeated token frequency.
include_reasoningbooleanSupportedIncludes reasoning content in the response when available.
logit_biasobjectSupportedAdjusts the likelihood of selected tokens.
logprobsbooleanSupportedReturns token log probabilities.
max_tokensintegerSupportedLimits generated output tokens.
min_pnumberSupportedApplies minimum-probability sampling.
presence_penaltynumberSupportedPenalizes tokens already present in the output.
reasoningobjectSupportedControls reasoning behavior and token allocation.
reasoning_effortenumSupportedSelects the model reasoning effort.
repetition_penaltynumberSupportedControls repetition across generated tokens.
response_formatobjectSupportedRequests a specific response format.
seedintegerSupportedRequests deterministic sampling when supported by the provider.
stopstring or arraySupportedStops generation at the supplied sequence.
structured_outputsbooleanSupportedEnables schema-constrained structured output.
temperaturenumberSupportedControls sampling randomness.
tool_choicestring or objectSupportedControls which tool the model may call.
toolsarraySupportedDefines tools available to the model.
top_avalueSupportedSupported by this model according to OpenRouter.
top_kintegerSupportedRestricts sampling to the highest-probability tokens.
top_logprobsintegerSupportedSets how many top-token log probabilities are returned.
top_pnumberSupportedControls nucleus sampling.

DeepSeek V4 Flash API Q&A

Model-specific answers for teams comparing DeepSeek V4 Flash API on capability, benchmark evidence, integration, and cost.

What is the best DeepSeek V4 Flash API benchmark rank?

The max configuration ranks #20 with a 36.24 median score; the standard configuration ranks #36.

How does DeepSeek V4 Flash API differ from DeepSeek V4 Pro?

Flash has substantially lower listed token prices and a lower current benchmark result, making it better suited to volume-sensitive routing.

Does DeepSeek V4 Flash API support image input?

No. The listed modality is text input to text output with a context window of roughly 1.05M tokens.

What is the DeepSeek V4 Flash cache price?

See the live pricing table above for the current API Need input and output prices for deepseek-v4-flash.

Which endpoint serves DeepSeek V4 Flash API?

Call DeepSeek V4 Flash API through https://apineed.com/v1/chat/completions with deepseek-v4-flash as the model value. Existing OpenAI SDK clients usually need only the APINEED base URL and key.

How do I validate DeepSeek V4 Flash API before release?

Evaluate DeepSeek V4 Flash API on representative prompts, record quality, latency, and token use, then choose reasoning settings and fallbacks from those results.

How to Deploy the DeepSeek V4 Flash API on apineed.com

Deploy DeepSeek V4 Flash API through APINEED after validating its model-specific trade-offs above. The API key, credit, and endpoint flow stays consistent across the catalog.

1

Add credits for DeepSeek V4 Flash API

Add credits on APINEED before deploying the DeepSeek V4 Flash API. Pay-as-you-go billing lets usage start small and scale with production traffic.

2

Get your API key

Create an APINEED API key for the DeepSeek V4 Flash API. The same key can call DeepSeek V4 Flash and other AI APIs through apineed.com.

3

Set the model slug

Use deepseek-v4-flash as the model value when you deploy the DeepSeek V4 Flash API. Keep the APINEED base URL at https://apineed.com/v1.

4

Send a request to DeepSeek V4 Flash API

Send chat completions or responses to the DeepSeek V4 Flash API from Claude Code, Codex, or any custom agent. The request keeps its OpenAI-compatible shape.

More models from DeepSeek