Plain-text specification
Plain-text specification
Google Gemma 4 26B A4B Instruct API
Google Gemma 4 26B A4B Instruct is a large language model by Google, available on the Venice API asgoogle-gemma-4-26b-a4b-it. It runs privately, with zero data retention.Gemma 4 26B A4B is a Mixture-of-Experts model from Google DeepMind with 26B total parameters and only 4B active per token, offering fast inference at high quality. It handles text, image, and video input, supports 256K context, function calling, and reasoning with configurable thinking modes.Google Gemma 4 26B A4B Instruct API pricing
Google Gemma 4 26B A4B Instruct specifications
How to use the Google Gemma 4 26B A4B Instruct API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "google-gemma-4-26b-a4b-it" and your API key.Google Gemma 4 26B A4B Instruct API FAQ
How much does the Google Gemma 4 26B A4B Instruct API cost?
0.40 per 1M output tokens, with cached input at $0.05 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Google Gemma 4 26B A4B Instruct model ID?
Usegoogle-gemma-4-26b-a4b-it as the model parameter.Is the Google Gemma 4 26B A4B Instruct API private?
Google Gemma 4 26B A4B Instruct is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Google Gemma 4 26B A4B Instruct?
250K tokens of context, with up to 8K output tokens per response.What does Google Gemma 4 26B A4B Instruct support?
Google Gemma 4 26B A4B Instruct supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default low).Which endpoint does the Google Gemma 4 26B A4B Instruct API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- DeepSeek V4 Flash 0423 API: 0.28 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- GLM 4.7 Flash Heretic API: 0.40 output per 1M tokens
- MiMo-V2.6-Flash API: 0.35 output per 1M tokens