Plain-text specification
Plain-text specification
Gemini 3.5 Flash-Lite API
Gemini 3.5 Flash-Lite is a large language model by Google, available on the Venice API asgemini-3-5-flash-lite. Requests are anonymized, so the provider never sees your identity.Gemini 3.5 Flash-Lite is the fastest, most cost-efficient Gemini 3.5 model with 1M context, ideal for everyday questions, summarization, and lightweight coding tasks.Gemini 3.5 Flash-Lite API pricing
Gemini 3.5 Flash-Lite specifications
How to use the Gemini 3.5 Flash-Lite API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "gemini-3-5-flash-lite" and your API key.Gemini 3.5 Flash-Lite API FAQ
How much does the Gemini 3.5 Flash-Lite API cost?
3.13 per 1M output tokens, with cached input at $0.037 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Gemini 3.5 Flash-Lite model ID?
Usegemini-3-5-flash-lite as the model parameter.Is the Gemini 3.5 Flash-Lite API private?
Gemini 3.5 Flash-Lite is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of Gemini 3.5 Flash-Lite?
1M tokens of context, with up to 64K output tokens per response.What does Gemini 3.5 Flash-Lite support?
Gemini 3.5 Flash-Lite supports function calling, structured outputs, reasoning, image input, web search and prompt caching.Which endpoint does the Gemini 3.5 Flash-Lite API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Aion 3.0 Mini API: 1.75 output per 1M tokens
- Qwen 3.6 27B API: 3.25 output per 1M tokens
- Qwen 3.8 27B API: 3.20 output per 1M tokens
- Aion 3.5 Mini API: 1.75 output per 1M tokens