Google: Gemini 3.5 Flash Lite API
Gemini 3.5 Flash Lite API minimizes input cost while retaining text, image, file, audio, and video support across a 1.05M context. It is aimed at broad first-pass processing rather than the highest logic score. [1]
Its high-reasoning row ranks #24 with a 22.82 median and 30.23 best score. The 188-second measured time is close to Gemini 3.5 Flash high, but the median is substantially lower. [1][2]
Playground
Test Gemini 3.5 Flash Lite API with representative production prompts before using gemini-3.5-flash-lite in a live workflow.
Providers
OpenRouter lists Gemini 3.5 Flash Lite API with text, image, video, file, and audio input, text output, and a 1.048576M context window. APINEED routing and prepaid rates are identified separately. [1]
Discount
The 50% APINEED discount changes official rates of $0.30 input, $2.50 output, and $0.03 cache read to $0.15, $1.25, and $0.015. The table compares official Gemini 3.5 Flash Lite API pricing with the APINEED prepaid rate.
Availability
APINEED continuously monitors Gemini 3.5 Flash Lite API access and keeps requests on healthy capacity.
Gemini 3.5 Flash Lite API Benchmarks
Gemini 3.5 Flash Lite high ranks #24. Its result sits below Gemini 3.6 Flash high, clarifying that Lite prioritizes cost rather than equivalent reasoning quality. The complete LLM2014 logic 2026-07 table remains below, with Gemini 3.5 Flash Lite API highlighted when a current row exists. [2]
| 1st | GPT-5.5 (xhigh) | 83.80 | 77.46 | 7.57% | 494s | 30,811 | $25.51 | $29.57 |
| 2nd | Kimi-K3 (max) | 82.91 | 74.80 | 9.78% | 1095s | 39,912 | $16.76 | $15.00 |
| 3rd | Claude Opus 4.8 (xhigh) | 82.62 | 66.70 | 19.27% | 791s | 41,833 | $28.86 | $24.64 |
| 4th | GPT-5.6 Sol (xhigh) | 81.94 | 72.99 | 10.92% | 369s | 17,433 | $14.43 | $29.57 |
| 5th | Claude Opus 5 (xhigh) | 78.38 | 71.88 | 8.29% | 426s | 26,392 | $18.21 | $24.64 |
| 6th | Qwen3.7-Max (xhigh) | 74.56 | 66.95 | 10.21% | 374s | 54,207 | $7.81 | $5.14 |
| 7th | GLM-5.2 (max) | 73.68 | 58.27 | 20.91% | 890s | 45,683 | $5.12 | $4.00 |
| 8th | Gemini 3.1 Pro (high) | 73.36 | 58.46 | 20.31% | 214s | 25,905 | $8.58 | $11.83 |
| 9th | GPT-5.6 Luna (xhigh) | 69.49 | 51.97 | 25.21% | 242s | 38,083 | $6.31 | $5.91 |
| 10th | DeepSeek V4 Pro (max) | 68.00 | 49.97 | 26.51% | 1468s | 60,558 | $1.45 | $0.86 |
| 11th | Grok 4.5 (high) | 67.59 | 56.72 | 16.08% | 451s | 43,372 | $7.18 | $5.91 |
| 12th | Muse Spark 1.1 | 66.18 | 57.23 | 13.52% | 542s | 42,416 | $4.98 | $4.19 |
| 13th | Gemini 3.5 Flash (high) | 65.92 | 60.39 | 8.39% | 188s | 36,372 | $9.03 | $8.87 |
| 14th | Doubao-Seed-2.1-pro (high) | 58.69 | 49.35 | 15.91% | 2003s | 85,059 | $10.21 | $4.29 |
| 15th | Qwen3.7-Plus (high) | 58.50 | 44.09 | 24.63% | 1237s | 57,152 | $1.83 | $1.14 |
| 16th | Gemini 3.6 Flash (high) | 56.87 | 41.74 | 26.6% | 206s | 26,329 | $5.45 | $7.39 |
| 17th | Tencent Hy3 (high) | 54.42 | 44.60 | 18.04% | 1035s | 52,037 | $0.83 | $0.57 |
| 18th | Claude Sonnet 5 (xhigh) | 51.37 | 43.14 | 16.02% | 529s | 44,118 | $12.18 | $9.86 |
| 19th | DeepSeek V4 Flash (max) | 50.62 | 36.24 | 28.41% | 611s | 49,593 | $0.40 | $0.29 |
| 20th | MiniMax-M3 | 49.18 | 40.94 | 16.75% | 985s | 56,467 | $1.90 | $1.20 |
| 21st | Claude Opus 5 | 42.08 | 31.73 | 24.6% | 154s | 10,180 | $7.02 | $24.64 |
| 22nd | Claude Opus 4.6 | 36.88 | 25.79 | 30.07% | 65s | 4,387 | $3.03 | $24.64 |
| 23rd | Doubao-Seed-2.0-lite 0428 (high) | 35.32 | 26.77 | 24.21% | 656s | 26,726 | $0.38 | $0.51 |
| 24th | Gemini 3.5 Flash (minimal) | 33.24 | 22.22 | 33.15% | 44s | 6,189 | $1.54 | $8.87 |
| 25th | Ling-3.0-flash | 32.53 | 20.62 | 36.61% | 688s | 86,204 | $0.00 | $0.00 |
| 26th | Gemini 3.5 Flash Lite (high) | 30.23 | 22.82 | 24.51% | 188s | 18,297 | $1.26 | $2.46 |
| 27th | GPT-5.5 Instant | 28.87 | 17.66 | 38.83% | 25s | 1,673 | $1.39 | $29.57 |
| 28th | Gemma 4 31B | 27.91 | 22.04 | 21.03% | 770s | 15,050 | $0.17 | $0.39 |
| 29th | MiMo-V2.5-Pro | 26.91 | 13.00 | 51.69% | 478s | 29,443 | $0.71 | $0.86 |
| 30th | Qwen3.7-Plus | 26.19 | 16.23 | 38.03% | 314s | 8,692 | $0.28 | $1.14 |
| 31st | Step-3.7-Flash | 26.18 | 13.74 | 47.52% | 258s | 44,912 | $1.46 | $1.16 |
| 32nd | Qwen3.5-27B | 24.96 | 17.98 | 27.96% | 451s | 26,391 | $0.51 | $0.69 |
| 33rd | Gemini 3.1 Flash Lite (high) | 23.78 | 14.30 | 39.87% | 82s | 28,351 | $1.17 | $1.48 |
| 34th | Claude Sonnet 5 | 23.64 | 10.86 | 54.06% | 109s | 5,962 | $1.65 | $9.86 |
| 35th | Qwen3.7-Max | 22.45 | 18.53 | 17.46% | 196s | 6,710 | $0.97 | $5.14 |
| 36th | ERNIE 5.1 | 20.49 | 15.49 | 24.4% | 457s | 24,093 | $1.73 | $2.57 |
| 37th | openPangu-2.0-Flash | 19.83 | 10.76 | 45.74% | 571s | 28,199 | $0.18 | $0.23 |
| 38th | LongCat-2.0 | 19.40 | 9.74 | 49.79% | 395s | 14,886 | $0.48 | $1.14 |
| 39th | DeepSeek V4 Flash | 19.39 | 13.37 | 31.05% | 63s | 5,286 | $0.04 | $0.29 |
| 40th | Mistral Medium 3.5 | 18.44 | 13.87 | 24.78% | 194s | 27,854 | $5.77 | $7.39 |
| 41st | Doubao-Seed-2.1-pro | 17.49 | 9.74 | 44.31% | 305s | 6,009 | $0.72 | $4.29 |
| 42nd | GLM-5.2 | 14.80 | 10.43 | 29.53% | 23s | 953 | $0.11 | $4.00 |
| 43rd | Ling-2.6-1T | 13.05 | 8.39 | 35.71% | 229s | 4,770 | $0.31 | $2.29 |
| 44th | iFLYTEK Spark X2 | 12.39 | 6.21 | 49.88% | 458s | 11,285 | $0.09 | $0.29 |
| 45th | Gemini 3.1 Flash Lite | 11.79 | 10.28 | 12.81% | 9s | 773 | $0.03 | $1.48 |
| 46th | Mistral Medium 3.5 | 10.83 | 7.13 | 34.16% | 17s | 2,546 | $0.53 | $7.39 |
| 47th | Ling-2.6-flash | 10.07 | 4.85 | 51.84% | 25s | 4,997 | $0.04 | $0.30 |
Quick Start
Connect Gemini 3.5 Flash Lite API without changing the OpenAI-style request shape. Use it for extraction, classification, media triage, and routing before escalating difficult cases. Structured outputs make its low input rate useful in repeatable pipelines.
Get your API key
Create an APINEED key for Gemini 3.5 Flash Lite API and keep it in an environment variable.
export API_NEED_API_KEY=sk-apineed-v1-...Make your first request
Use gemini-3.5-flash-lite for Gemini 3.5 Flash Lite API with the APINEED API. The request shape is compatible with OpenAI chat completions, so most SDKs only need a base URL change.
import OpenAI from "openai";
const client = new OpenAI({
apiKey: process.env.API_NEED_API_KEY,
baseURL: "https://apineed.com/v1"
});
const result = await client.chat.completions.create({
model: "gemini-3.5-flash-lite",
messages: [
{ role: "user", content: "Why is the sky blue?" }
]
});
console.log(result.choices[0].message.content);Enable streaming and fallbacks
Add stream: true when Gemini 3.5 Flash Lite API should return server-sent events. APINEED keeps routing, provider health, and fallback handling behind the same endpoint.
curl https://apineed.com/v1/chat/completions \
-H "Authorization: Bearer $API_NEED_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gemini-3.5-flash-lite",
"stream": true,
"messages": [{ "role": "user", "content": "Hello" }]
}'Gemini 3.5 Flash Lite API Endpoint
Gemini 3.5 Flash Lite API accepts chat conversations here with streaming or non-streaming text output; the supported controls are listed on this page.
/v1/chat/completions- Authorization
- Bearer $API_NEED_API_KEY
- Content-Type
- application/json
- HTTP-Referer
- optional - your site URL, for rankings
- X-Title
- optional - your site name, for rankings
- Model
- gemini-3.5-flash-lite
Parameters
Gemini 3.5 Flash Lite API parameters currently listed for gemini-3.5-flash-lite by the OpenRouter Models API. [1]
| Name | Type | Status | Description |
|---|---|---|---|
include_reasoning | boolean | Supported | Includes reasoning content in the response when available. |
max_tokens | integer | Supported | Limits generated output tokens. |
reasoning | object | Supported | Controls reasoning behavior and token allocation. |
reasoning_effort | enum | Supported | Selects the model reasoning effort. |
response_format | object | Supported | Requests a specific response format. |
seed | integer | Supported | Requests deterministic sampling when supported by the provider. |
stop | string or array | Supported | Stops generation at the supplied sequence. |
structured_outputs | boolean | Supported | Enables schema-constrained structured output. |
temperature | number | Supported | Controls sampling randomness. |
tool_choice | string or object | Supported | Controls which tool the model may call. |
tools | array | Supported | Defines tools available to the model. |
top_p | number | Supported | Controls nucleus sampling. |
Gemini 3.5 Flash Lite API Q&A
Model-specific answers for teams comparing Gemini 3.5 Flash Lite API on capability, benchmark evidence, integration, and cost.
What makes Gemini 3.5 Flash Lite API economical?
Its APINEED input rate is $0.15 per million tokens and cache reads cost $0.015 after the 50% discount.
Can Gemini 3.5 Flash Lite API process video and audio?
Yes. Text, image, file, audio, and video are listed input modalities, with text output.
Where does Gemini 3.5 Flash Lite API rank?
The high configuration ranks #24 with a 22.82 median in the current logic dataset.
Which workloads fit Gemini 3.5 Flash Lite?
For Gemini 3.5 Flash Lite API, high-volume extraction, tagging, summarization, and multimodal routing are stronger fits than the most difficult reasoning tasks.
Which endpoint serves Gemini 3.5 Flash Lite API?
Call Gemini 3.5 Flash Lite API through https://apineed.com/v1/chat/completions with gemini-3.5-flash-lite as the model value. Existing OpenAI SDK clients usually need only the APINEED base URL and key.
How do I validate Gemini 3.5 Flash Lite API before release?
Evaluate Gemini 3.5 Flash Lite API on representative prompts, record quality, latency, and token use, then choose reasoning settings and fallbacks from those results.
How to Deploy the Gemini 3.5 Flash Lite API on apineed.com
Deploy Gemini 3.5 Flash Lite API through APINEED after validating its model-specific trade-offs above. The API key, credit, and endpoint flow stays consistent across the catalog.
Add credits for Gemini 3.5 Flash Lite API
Add credits on APINEED before deploying the Gemini 3.5 Flash Lite API. Pay-as-you-go billing lets usage start small and scale with production traffic.
Get your API key
Create an APINEED API key for the Gemini 3.5 Flash Lite API. The same key can call Gemini 3.5 Flash Lite and other AI APIs through apineed.com.
Set the model slug
Use gemini-3.5-flash-lite as the model value when you deploy the Gemini 3.5 Flash Lite API. Keep the APINEED base URL at https://apineed.com/v1.
Send a request to Gemini 3.5 Flash Lite API
Send chat completions or responses to the Gemini 3.5 Flash Lite API from Claude Code, Codex, or any custom agent. The request keeps its OpenAI-compatible shape.
