Skip to main content

Gemini 3.5 Flash-Lite API

Gemini 3.5 Flash-Lite is a large language model by Google, available on the Venice API as gemini-3-5-flash-lite. Requests are anonymized, so the provider never sees your identity.Gemini 3.5 Flash-Lite is the fastest, most cost-efficient Gemini 3.5 model with 1M context, ideal for everyday questions, summarization, and lightweight coding tasks.

Gemini 3.5 Flash-Lite API pricing

Gemini 3.5 Flash-Lite specifications

How to use the Gemini 3.5 Flash-Lite API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "gemini-3-5-flash-lite" and your API key.

Gemini 3.5 Flash-Lite API FAQ

How much does the Gemini 3.5 Flash-Lite API cost?

0.38per1Minputtokensand0.38 per 1M input tokens and 3.13 per 1M output tokens, with cached input at $0.037 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Gemini 3.5 Flash-Lite model ID?

Use gemini-3-5-flash-lite as the model parameter.

Is the Gemini 3.5 Flash-Lite API private?

Gemini 3.5 Flash-Lite is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of Gemini 3.5 Flash-Lite?

1M tokens of context, with up to 64K output tokens per response.

What does Gemini 3.5 Flash-Lite support?

Gemini 3.5 Flash-Lite supports function calling, structured outputs, reasoning, image input, web search and prompt caching.

Which endpoint does the Gemini 3.5 Flash-Lite API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models