Skip to main content

Google Gemma 4 31B Instruct API

Google Gemma 4 31B Instruct is a large language model by Google, available on the Venice API as google-gemma-4-31b-it. It runs privately, with zero data retention.Gemma 4 31B is a dense model from Google DeepMind with 31B parameters, delivering frontier-level reasoning performance. It handles text, image, and video input, supports 256K context, function calling, and configurable thinking modes.

Google Gemma 4 31B Instruct API pricing

Google Gemma 4 31B Instruct specifications

How to use the Google Gemma 4 31B Instruct API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "google-gemma-4-31b-it" and your API key.

Google Gemma 4 31B Instruct API FAQ

How much does the Google Gemma 4 31B Instruct API cost?

0.12per1Minputtokensand0.12 per 1M input tokens and 0.36 per 1M output tokens, with cached input at $0.09 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Google Gemma 4 31B Instruct model ID?

Use google-gemma-4-31b-it as the model parameter.

Is the Google Gemma 4 31B Instruct API private?

Google Gemma 4 31B Instruct is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Google Gemma 4 31B Instruct?

250K tokens of context, with up to 8K output tokens per response.

What does Google Gemma 4 31B Instruct support?

Google Gemma 4 31B Instruct supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default low).

Which endpoint does the Google Gemma 4 31B Instruct API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models