Google: Gemini 3.5 Flash API
Gemini 3.5 Flash API supports text, images, files, audio, and video with a 1.05M context, positioning it for complex multimodal agents. Its high mode is one of the strongest current Flash results. [1]
High mode ranks #7 with a 60.39 median, while minimal mode ranks #25 with a 22.22 median and completes 144 seconds faster. Reasoning mode changes both quality and interaction speed. [1][2]
Playground
Test Gemini 3.5 Flash API with representative production prompts before using gemini-3.5-flash in a live workflow.
Providers
OpenRouter lists Gemini 3.5 Flash API with text, image, video, file, and audio input, text output, and a 1.048576M context window. APINEED routing and prepaid rates are identified separately. [1]
Discount
APINEED halves official rates to $0.75 input, $4.50 output, $0.075 cache read, and about $0.04167 cache create per million tokens. The table compares official Gemini 3.5 Flash API pricing with the APINEED prepaid rate.
Availability
APINEED continuously monitors Gemini 3.5 Flash API access and keeps requests on healthy capacity.
Gemini 3.5 Flash API Benchmarks
Gemini 3.5 Flash high ranks #7, immediately ahead of Gemini 3.1 Pro high. Minimal mode ranks #25 and records a much shorter measured completion time. The complete LLM2014 logic 2026-07 table remains below, with Gemini 3.5 Flash API highlighted when a current row exists. [2]
| 1st | GPT-5.5 (xhigh) | 83.80 | 77.46 | 7.57% | 494s | 30,811 | $25.51 | $29.57 |
| 2nd | Kimi-K3 (max) | 82.91 | 74.80 | 9.78% | 1095s | 39,912 | $16.76 | $15.00 |
| 3rd | Claude Opus 4.8 (xhigh) | 82.62 | 66.70 | 19.27% | 791s | 41,833 | $28.86 | $24.64 |
| 4th | GPT-5.6 Sol (xhigh) | 81.94 | 72.99 | 10.92% | 369s | 17,433 | $14.43 | $29.57 |
| 5th | Claude Opus 5 (xhigh) | 78.38 | 71.88 | 8.29% | 426s | 26,392 | $18.21 | $24.64 |
| 6th | Qwen3.7-Max (xhigh) | 74.56 | 66.95 | 10.21% | 374s | 54,207 | $7.81 | $5.14 |
| 7th | GLM-5.2 (max) | 73.68 | 58.27 | 20.91% | 890s | 45,683 | $5.12 | $4.00 |
| 8th | Gemini 3.1 Pro (high) | 73.36 | 58.46 | 20.31% | 214s | 25,905 | $8.58 | $11.83 |
| 9th | GPT-5.6 Luna (xhigh) | 69.49 | 51.97 | 25.21% | 242s | 38,083 | $6.31 | $5.91 |
| 10th | DeepSeek V4 Pro (max) | 68.00 | 49.97 | 26.51% | 1468s | 60,558 | $1.45 | $0.86 |
| 11th | Grok 4.5 (high) | 67.59 | 56.72 | 16.08% | 451s | 43,372 | $7.18 | $5.91 |
| 12th | Muse Spark 1.1 | 66.18 | 57.23 | 13.52% | 542s | 42,416 | $4.98 | $4.19 |
| 13th | Gemini 3.5 Flash (high) | 65.92 | 60.39 | 8.39% | 188s | 36,372 | $9.03 | $8.87 |
| 14th | Doubao-Seed-2.1-pro (high) | 58.69 | 49.35 | 15.91% | 2003s | 85,059 | $10.21 | $4.29 |
| 15th | Qwen3.7-Plus (high) | 58.50 | 44.09 | 24.63% | 1237s | 57,152 | $1.83 | $1.14 |
| 16th | Gemini 3.6 Flash (high) | 56.87 | 41.74 | 26.6% | 206s | 26,329 | $5.45 | $7.39 |
| 17th | Tencent Hy3 (high) | 54.42 | 44.60 | 18.04% | 1035s | 52,037 | $0.83 | $0.57 |
| 18th | Claude Sonnet 5 (xhigh) | 51.37 | 43.14 | 16.02% | 529s | 44,118 | $12.18 | $9.86 |
| 19th | DeepSeek V4 Flash (max) | 50.62 | 36.24 | 28.41% | 611s | 49,593 | $0.40 | $0.29 |
| 20th | MiniMax-M3 | 49.18 | 40.94 | 16.75% | 985s | 56,467 | $1.90 | $1.20 |
| 21st | Claude Opus 5 | 42.08 | 31.73 | 24.6% | 154s | 10,180 | $7.02 | $24.64 |
| 22nd | Claude Opus 4.6 | 36.88 | 25.79 | 30.07% | 65s | 4,387 | $3.03 | $24.64 |
| 23rd | Doubao-Seed-2.0-lite 0428 (high) | 35.32 | 26.77 | 24.21% | 656s | 26,726 | $0.38 | $0.51 |
| 24th | Gemini 3.5 Flash (minimal) | 33.24 | 22.22 | 33.15% | 44s | 6,189 | $1.54 | $8.87 |
| 25th | Ling-3.0-flash | 32.53 | 20.62 | 36.61% | 688s | 86,204 | $0.00 | $0.00 |
| 26th | Gemini 3.5 Flash Lite (high) | 30.23 | 22.82 | 24.51% | 188s | 18,297 | $1.26 | $2.46 |
| 27th | GPT-5.5 Instant | 28.87 | 17.66 | 38.83% | 25s | 1,673 | $1.39 | $29.57 |
| 28th | Gemma 4 31B | 27.91 | 22.04 | 21.03% | 770s | 15,050 | $0.17 | $0.39 |
| 29th | MiMo-V2.5-Pro | 26.91 | 13.00 | 51.69% | 478s | 29,443 | $0.71 | $0.86 |
| 30th | Qwen3.7-Plus | 26.19 | 16.23 | 38.03% | 314s | 8,692 | $0.28 | $1.14 |
| 31st | Step-3.7-Flash | 26.18 | 13.74 | 47.52% | 258s | 44,912 | $1.46 | $1.16 |
| 32nd | Qwen3.5-27B | 24.96 | 17.98 | 27.96% | 451s | 26,391 | $0.51 | $0.69 |
| 33rd | Gemini 3.1 Flash Lite (high) | 23.78 | 14.30 | 39.87% | 82s | 28,351 | $1.17 | $1.48 |
| 34th | Claude Sonnet 5 | 23.64 | 10.86 | 54.06% | 109s | 5,962 | $1.65 | $9.86 |
| 35th | Qwen3.7-Max | 22.45 | 18.53 | 17.46% | 196s | 6,710 | $0.97 | $5.14 |
| 36th | ERNIE 5.1 | 20.49 | 15.49 | 24.4% | 457s | 24,093 | $1.73 | $2.57 |
| 37th | openPangu-2.0-Flash | 19.83 | 10.76 | 45.74% | 571s | 28,199 | $0.18 | $0.23 |
| 38th | LongCat-2.0 | 19.40 | 9.74 | 49.79% | 395s | 14,886 | $0.48 | $1.14 |
| 39th | DeepSeek V4 Flash | 19.39 | 13.37 | 31.05% | 63s | 5,286 | $0.04 | $0.29 |
| 40th | Mistral Medium 3.5 | 18.44 | 13.87 | 24.78% | 194s | 27,854 | $5.77 | $7.39 |
| 41st | Doubao-Seed-2.1-pro | 17.49 | 9.74 | 44.31% | 305s | 6,009 | $0.72 | $4.29 |
| 42nd | GLM-5.2 | 14.80 | 10.43 | 29.53% | 23s | 953 | $0.11 | $4.00 |
| 43rd | Ling-2.6-1T | 13.05 | 8.39 | 35.71% | 229s | 4,770 | $0.31 | $2.29 |
| 44th | iFLYTEK Spark X2 | 12.39 | 6.21 | 49.88% | 458s | 11,285 | $0.09 | $0.29 |
| 45th | Gemini 3.1 Flash Lite | 11.79 | 10.28 | 12.81% | 9s | 773 | $0.03 | $1.48 |
| 46th | Mistral Medium 3.5 | 10.83 | 7.13 | 34.16% | 17s | 2,546 | $0.53 | $7.39 |
| 47th | Ling-2.6-flash | 10.07 | 4.85 | 51.84% | 25s | 4,997 | $0.04 | $0.30 |
Quick Start
Connect Gemini 3.5 Flash API without changing the OpenAI-style request shape. Use high mode for cross-media reasoning and minimal mode for quicker extraction or routing. Tools and structured outputs allow both modes to share one integration.
Get your API key
Create an APINEED key for Gemini 3.5 Flash API and keep it in an environment variable.
export API_NEED_API_KEY=sk-apineed-v1-...Make your first request
Use gemini-3.5-flash for Gemini 3.5 Flash API with the APINEED API. The request shape is compatible with OpenAI chat completions, so most SDKs only need a base URL change.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.API_NEED_API_KEY,
baseURL: "https://apineed.com/v1"
});
const result = await client.chat.completions.create({
model: "gemini-3.5-flash",
messages: [
{ role: "user", content: "Why is the sky blue?" }
]
});
console.log(result.choices[0].message.content);Enable streaming and fallbacks
Add stream: true when Gemini 3.5 Flash API should return server-sent events. APINEED keeps routing, provider health, and fallback handling behind the same endpoint.
curl https://apineed.com/v1/chat/completions \
-H "Authorization: Bearer $API_NEED_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.5-flash",
"stream": true,
"messages": [{ "role": "user", "content": "Hello" }]
}'Gemini 3.5 Flash API Endpoint
Gemini 3.5 Flash API accepts chat conversations here with streaming or non-streaming text output; the supported controls are listed on this page.
/v1/chat/completions- Authorization
- Bearer $API_NEED_API_KEY
- Content-Type
- application/json
- HTTP-Referer
- optional - your site URL, for rankings
- X-Title
- optional - your site name, for rankings
- Model
- gemini-3.5-flash
Parameters
Gemini 3.5 Flash API parameters currently listed for gemini-3.5-flash by the OpenRouter Models API. [1]
| Name | Type | Status | Description |
|---|---|---|---|
include_reasoning | boolean | Supported | Includes reasoning content in the response when available. |
max_tokens | integer | Supported | Limits generated output tokens. |
reasoning | object | Supported | Controls reasoning behavior and token allocation. |
reasoning_effort | enum | Supported | Selects the model reasoning effort. |
response_format | object | Supported | Requests a specific response format. |
seed | integer | Supported | Requests deterministic sampling when supported by the provider. |
stop | string or array | Supported | Stops generation at the supplied sequence. |
structured_outputs | boolean | Supported | Enables schema-constrained structured output. |
temperature | number | Supported | Controls sampling randomness. |
tool_choice | string or object | Supported | Controls which tool the model may call. |
tools | array | Supported | Defines tools available to the model. |
top_p | number | Supported | Controls nucleus sampling. |
Gemini 3.5 Flash API Q&A
Model-specific answers for teams comparing Gemini 3.5 Flash API on capability, benchmark evidence, integration, and cost.
Which Gemini 3.5 Flash API mode ranks highest?
High mode ranks #7 with a 60.39 median; minimal mode ranks #25 with a 22.22 median.
Does Gemini 3.5 Flash API accept audio and video?
Yes. It accepts text, images, files, audio, and video and produces text output.
What is the Gemini 3.5 Flash APINEED output price?
See the live pricing table above for the current API Need input and output prices for gemini-3.5-flash.
When is minimal mode useful for Gemini 3.5 Flash?
For Gemini 3.5 Flash API, minimal mode is appropriate when lower latency matters more than high-mode logic performance for routine agent steps.
Which endpoint serves Gemini 3.5 Flash API?
Call Gemini 3.5 Flash API through https://apineed.com/v1/chat/completions with gemini-3.5-flash as the model value. Existing OpenAI SDK clients usually need only the APINEED base URL and key.
How do I validate Gemini 3.5 Flash API before release?
Evaluate Gemini 3.5 Flash API on representative prompts, record quality, latency, and token use, then choose reasoning settings and fallbacks from those results.
How to Deploy the Gemini 3.5 Flash API on apineed.com
Deploy Gemini 3.5 Flash API through APINEED after validating its model-specific trade-offs above. The API key, credit, and endpoint flow stays consistent across the catalog.
Add credits for Gemini 3.5 Flash API
Add credits on APINEED before deploying the Gemini 3.5 Flash API. Pay-as-you-go billing lets usage start small and scale with production traffic.
Get your API key
Create an APINEED API key for the Gemini 3.5 Flash API. The same key can call Gemini 3.5 Flash and other AI APIs through apineed.com.
Set the model slug
Use gemini-3.5-flash as the model value when you deploy the Gemini 3.5 Flash API. Keep the APINEED base URL at https://apineed.com/v1.
Send a request to Gemini 3.5 Flash API
Send chat completions or responses to the Gemini 3.5 Flash API from Claude Code, Codex, or any custom agent. The request keeps its OpenAI-compatible shape.
