Skip to main content

Grok 4.6 API

Grok 4.6 is a large language model by xAI, available on the Venice API as grok-4-6. It runs privately, with zero data retention.Grok 4.6 is xAI’s multimodal chat and reasoning model with function calling, structured outputs, adjustable reasoning effort (low/medium/high/xhigh), and a 500K-token context window.

Grok 4.6 API pricing

Grok 4.6 specifications

How to use the Grok 4.6 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "grok-4-6" and your API key.

Grok 4.6 API FAQ

How much does the Grok 4.6 API cost?

2.27per1Minputtokensand2.27 per 1M input tokens and 6.80 per 1M output tokens, with cached input at $0.57 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Grok 4.6 model ID?

Use grok-4-6 as the model parameter.

Is the Grok 4.6 API private?

Grok 4.6 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Grok 4.6?

500K tokens of context, with up to 200K output tokens per response.

What does Grok 4.6 support?

Grok 4.6 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: low, medium, high and xhigh (default high).

Which endpoint does the Grok 4.6 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models

  • Grok 4.7 API: 2.27input/2.27 input / 6.80 output per 1M tokens
  • Grok 4.5 API: 2.27input/2.27 input / 6.80 output per 1M tokens
  • Grok 4.3 API: 1.42input/1.42 input / 2.83 output per 1M tokens
  • Grok 4.20 API: 1.42input/1.42 input / 2.83 output per 1M tokens
  • Qwen 3.8 2.4T API: 2.50input/2.50 input / 7.50 output per 1M tokens
  • Qwen 3.8 Max API: 2.50input/2.50 input / 7.50 output per 1M tokens
  • Gemini 3.5 Flash API: 1.55input/1.55 input / 9.45 output per 1M tokens
  • GLM 5.3 API: 1.75input/1.75 input / 5.50 output per 1M tokens