Plain-text specification
Plain-text specification
Kimi K3 API
Kimi K3 is a large language model by Moonshot AI, available on the Venice API askimi-k3. Private and end-to-end encrypted variants are available.Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.Kimi K3 API pricing
Kimi K3 specifications
How to use the Kimi K3 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "kimi-k3" and your API key.Kimi K3 API FAQ
How much does the Kimi K3 API cost?
18.75 per 1M output tokens, with cached input at 4.50 input and $22.50 output. Prices are in USD and can be paid in DIEM at parity.What is the Kimi K3 model ID?
Usekimi-k3 as the model parameter. Other variants: kimi-k3-fast-api (Fast), e2ee-kimi-k3-p (E2EE).Is the Kimi K3 API private?
The Standard and Fast variants are private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.What is the context window of Kimi K3?
1M tokens of context, with up to 128K output tokens per response.What does Kimi K3 support?
Kimi K3 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, high and max (default max).Which endpoint does the Kimi K3 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Kimi K2.6 API: 3.50 output per 1M tokens
- Kimi K2.5 API: 3.50 output per 1M tokens
- Claude Sonnet 5.5 API: 18.75 output per 1M tokens
- GPT-5.4 API: 18.80 output per 1M tokens
- Claude Sonnet 4.6 API: 18.00 output per 1M tokens
- Claude Sonnet 5 API: 15.00 output per 1M tokens