Skip to main content

Kimi K2.6 API

Kimi K2.6 is a large language model by Moonshot AI, available on the Venice API as kimi-k2-6. Private and end-to-end encrypted variants are available.Kimi K2.6 is an open-source, native multimodal agentic model from Moonshot AI with 1T total parameters and 32B active parameters. It excels at long-horizon coding, coding-driven design, agent swarm orchestration, and proactive autonomous execution with 256K context windows.

Kimi K2.6 API pricing

Kimi K2.6 specifications

How to use the Kimi K2.6 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "kimi-k2-6" and your API key.

Kimi K2.6 API FAQ

How much does the Kimi K2.6 API cost?

0.75per1Minputtokensand0.75 per 1M input tokens and 3.50 per 1M output tokens, with cached input at 0.16per1M.TheE2EEvariantcosts0.16 per 1M. The E2EE variant costs 0.87 input and $4.12 output. Prices are in USD and can be paid in DIEM at parity.

What is the Kimi K2.6 model ID?

Use kimi-k2-6 as the model parameter. Other variants: e2ee-kimi-k2-6 (E2EE).

Is the Kimi K2.6 API private?

The Standard variant is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.

What is the context window of Kimi K2.6?

250K tokens of context, with up to 64K output tokens per response.

What does Kimi K2.6 support?

Kimi K2.6 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the Kimi K2.6 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models