> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# Kimi K3 API

> Kimi K3 API on Venice: 1M context, $3.75 input and $18.75 output per 1M tokens. Private and end-to-end encrypted variants. Pricing, specs and code examples.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"kimi-k3","name":"Kimi K3","modality":"text","task":"chat","provider":"moonshot","description":"Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.","primary":"kimi-k3","variants":["kimi-k3","kimi-k3-fast-api","e2ee-kimi-k3-p"],"created":1784160000,"updated":1788825600,"privacy":["private","e2ee"],"openWeights":true,"traits":["default_reasoning"],"sets":["venice_recommendations"]},"models":[{"id":"kimi-k3","name":"Kimi K3","type":"text","modality":"text","task":"chat","variant":"standard","provider":"moonshot","created":1784160000,"description":"Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.","source":"https://huggingface.co/moonshotai/Kimi-K3","privacy":"private","traits":["default_reasoning"],"sets":["venice_recommendations"],"openWeights":true,"text":{"context":1000000,"maxOutput":131072,"reasoning":{"supported":true,"effort":["none","low","high","max"],"defaultEffort":"max"},"caps":{"tools":true,"structured":true,"vision":true,"maxImages":10,"webSearch":true,"code":true}},"pricing":{"input":3.75,"output":18.75,"cacheRead":0.375,"blended":7.5},"headline":{"value":7.5,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha"}],"family":"kimi-k3"},{"id":"kimi-k3-fast-api","name":"Kimi K3 Fast","type":"text","modality":"text","task":"chat","variant":"fast","provider":"moonshot","created":1785715200,"description":"Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.","source":"https://huggingface.co/moonshotai/Kimi-K3","privacy":"private","openWeights":true,"text":{"context":1000000,"maxOutput":131072,"reasoning":{"supported":true,"effort":["none","low","high","max"],"defaultEffort":"max"},"caps":{"tools":true,"structured":true,"vision":true,"maxImages":10,"webSearch":true,"code":true}},"pricing":{"input":4.5,"output":22.5,"cacheRead":0.45,"blended":9},"headline":{"value":9,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha"}],"family":"kimi-k3"},{"id":"e2ee-kimi-k3-p","name":"Kimi K3","type":"text","modality":"text","task":"chat","variant":"e2ee","provider":"moonshot","created":1788825600,"description":"Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows.","source":"https://huggingface.co/moonshotai/Kimi-K3","privacy":"e2ee","openWeights":true,"text":{"context":1000000,"maxOutput":131072,"reasoning":{"supported":true,"effort":["none","max"],"defaultEffort":"max"},"caps":{"tools":true,"structured":true,"vision":true,"maxImages":10,"webSearch":true,"code":true,"e2ee":true,"tee":true}},"pricing":{"input":3.75,"output":18.75,"cacheRead":0.375,"blended":7.5},"headline":{"value":7.5,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha","note":"E2EE models require /chat/completions with E2EE headers."}],"family":"kimi-k3"}],"related":{"similar":[{"slug":"claude-sonnet-5-5","name":"Claude Sonnet 5.5","provider":"anthropic","modality":"text","privacy":["anonymized"],"created":1790380800,"headline":{"value":7.5,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"gpt-5-4","name":"GPT-5.4","provider":"openai","modality":"text","privacy":["anonymized"],"created":1772668800,"headline":{"value":7.0475,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"claude-sonnet-4-6","name":"Claude Sonnet 4.6","provider":"anthropic","modality":"text","privacy":["anonymized"],"created":1771286400,"headline":{"value":7.2,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"claude-sonnet-5","name":"Claude Sonnet 5","provider":"anthropic","modality":"text","privacy":["anonymized"],"created":1782691200,"headline":{"value":6,"unit":"per 1M tokens","basis":"blended"},"variants":1}],"versions":[{"slug":"kimi-k2-6","name":"Kimi K2.6","provider":"moonshot","modality":"text","privacy":["private","e2ee"],"created":1776643200,"headline":{"value":1.4375,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"kimi-k2-5","name":"Kimi K2.5","provider":"moonshot","modality":"text","privacy":["private"],"created":1769548800,"headline":{"value":1.295,"unit":"per 1M tokens","basis":"blended"},"variants":1}]},"providers":{"moonshot":{"slug":"moonshot","name":"Moonshot AI","logo":"/images/icons/models/kimi.svg"},"anthropic":{"slug":"anthropic","name":"Anthropic","logo":"/images/icons/models/opus.svg"},"openai":{"slug":"openai","name":"OpenAI","logo":"/images/icons/models/openai.svg"}},"faq":[{"q":"How much does the Kimi K3 API cost?","a":"$3.75 per 1M input tokens and $18.75 per 1M output tokens, with cached input at $0.38 per 1M. The Fast variant costs $4.50 input and $22.50 output. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the Kimi K3 model ID?","a":"Use `kimi-k3` as the `model` parameter. Other variants: `kimi-k3-fast-api` (Fast), `e2ee-kimi-k3-p` (E2EE)."},{"q":"Is the Kimi K3 API private?","a":"The Standard and Fast variants are private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them."},{"q":"What is the context window of Kimi K3?","a":"1M tokens of context, with up to 128K output tokens per response."},{"q":"What does Kimi K3 support?","a":"Kimi K3 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, high and max (default max)."},{"q":"Which endpoint does the Kimi K3 API use?","a":"Call `POST /chat/completions`. `/responses` (Alpha) is also supported."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # Kimi K3 API

    Kimi K3 is a large language model by Moonshot AI, available on the Venice API as `kimi-k3`. Private and end-to-end encrypted variants are available.

    Kimi K3 is an ultra-large-scale, open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.

    ## Kimi K3 API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `kimi-k3` | Standard | Private | $3.75 input / $18.75 output per 1M tokens |
    | `kimi-k3-fast-api` | Fast | Private | $4.50 input / $22.50 output per 1M tokens |
    | `e2ee-kimi-k3-p` | E2EE | End-to-end encrypted | $3.75 input / $18.75 output per 1M tokens |

    ## Kimi K3 specifications

    | Spec | Value |
    | - | - |
    | Provider | Moonshot AI |
    | Released | Jul 16, 2026 |
    | Privacy | Private, End-to-end encrypted |
    | Open weights | Yes |
    | Context window | 1M tokens |
    | Max output | 128K tokens |
    | Reasoning effort | none, low, high, max |

    ## How to use the Kimi K3 API

    Send requests to `POST https://api.venice.ai/api/v1/chat/completions` with `"model": "kimi-k3"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "kimi-k3",
        "messages": [{ "role": "user", "content": "Explain TEE attestation in two sentences." }],
        "reasoning_effort": "max"
      }'
    ```

    ## Kimi K3 API FAQ

    ### How much does the Kimi K3 API cost?

    $3.75 per 1M input tokens and $18.75 per 1M output tokens, with cached input at $0.38 per 1M. The Fast variant costs $4.50 input and \$22.50 output. Prices are in USD and can be paid in DIEM at parity.

    ### What is the Kimi K3 model ID?

    Use `kimi-k3` as the `model` parameter. Other variants: `kimi-k3-fast-api` (Fast), `e2ee-kimi-k3-p` (E2EE).

    ### Is the Kimi K3 API private?

    The Standard and Fast variants are private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.

    ### What is the context window of Kimi K3?

    1M tokens of context, with up to 128K output tokens per response.

    ### What does Kimi K3 support?

    Kimi K3 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: none, low, high and max (default max).

    ### Which endpoint does the Kimi K3 API use?

    Call `POST /chat/completions`. `/responses` (Alpha) is also supported.

    ## Related models

    * [Kimi K2.6 API](/models/kimi-k2-6): $0.75 input / $3.50 output per 1M tokens
    * [Kimi K2.5 API](/models/kimi-k2-5): $0.56 input / $3.50 output per 1M tokens
    * [Claude Sonnet 5.5 API](/models/claude-sonnet-5-5): $3.75 input / $18.75 output per 1M tokens
    * [GPT-5.4 API](/models/gpt-5-4): $3.13 input / $18.80 output per 1M tokens
    * [Claude Sonnet 4.6 API](/models/claude-sonnet-4-6): $3.60 input / $18.00 output per 1M tokens
    * [Claude Sonnet 5 API](/models/claude-sonnet-5): $3.00 input / $15.00 output per 1M tokens
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.