> ## Documentation Index
> Fetch the complete documentation index at: https://veniceai-feat-models-redesign.mintlify.site/llms.txt
> Use this file to discover all available pages before exploring further.

# GLM 4.6 API

> GLM 4.6 API on Venice: 198K context, $0.43 input and $1.75 output per 1M tokens. Private, with zero data retention. Supports reasoning and function calling.

export const HubMount = ({view, children, ...props}) => {
  const [hub, setHub] = useState(null);
  const [failed, setFailed] = useState(false);
  useEffect(() => {
    let alive = true;
    const w = window;
    if (!w.__veniceModelHub) {
      const urls = w.location.hostname === 'localhost' ? ['http://localhost:3333/data/model-hub.bundle.json', '/data/model-hub.bundle.json'] : ['/data/model-hub.bundle.json'];
      const load = i => fetch(urls[i], {
        cache: 'no-cache'
      }).then(res => {
        if (!res.ok) throw new Error(`bundle ${res.status}`);
        return res.json();
      }).catch(err => i + 1 < urls.length ? load(i + 1) : Promise.reject(err));
      const Frag = <></>.type;
      const h = (type, props, ...kids) => {
        const T = type;
        const {key, ...rest} = props || ({});
        if (!kids.length) return <T key={key} {...rest} />;
        if (kids.length === 1) return <T key={key} {...rest}>{kids[0]}</T>;
        return <T key={key} {...rest}>{kids.map((kid, i) => <Frag key={i}>{kid}</Frag>)}</T>;
      };
      w.__veniceModelHub = load(0).then(bundle => new Function(`return (${bundle.code})`)()({
        h,
        Fragment: Frag,
        useState,
        useEffect,
        useRef,
        useMemo,
        useCallback
      }));
    }
    w.__veniceModelHub.then(instance => {
      if (alive) setHub(instance);
    }).catch(() => {
      w.__veniceModelHub = null;
      if (alive) setFailed(true);
    });
    return () => {
      alive = false;
    };
  }, []);
  const View = hub ? hub[view] : null;
  if (View) return <View {...props}>{children}</View>;
  if (failed) {
    return <div className="vx-mount is-failed">
        <p className="vx-mount-note">The interactive model catalog could not load. The full data is below.</p>
        {children}
      </div>;
  }
  return <div className="vx-mount" aria-busy="true">
      <div className="vx-mount-skeleton" aria-hidden="true"><span /><span /><span /></div>
      <div className="vx-mount-source">{children}</div>
    </div>;
};

<HubMount view="ModelPage" data={{"family":{"slug":"glm-4-6","name":"GLM 4.6","modality":"text","task":"chat","provider":"zai","description":"GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.","primary":"zai-org-glm-4.6","variants":["zai-org-glm-4.6"],"created":1711929600,"updated":1711929600,"privacy":["private"],"openWeights":true},"models":[{"id":"zai-org-glm-4.6","name":"GLM 4.6","type":"text","modality":"text","task":"chat","variant":"standard","provider":"zai","created":1711929600,"description":"GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.","source":"https://huggingface.co/zai-org/GLM-4.6","privacy":"private","openWeights":true,"text":{"context":198000,"maxOutput":16384,"quantization":"fp4","reasoning":{"supported":true,"effort":["low","medium","high"],"defaultEffort":"low"},"caps":{"tools":true,"structured":true,"webSearch":true}},"pricing":{"input":0.43,"output":1.75,"cacheRead":0.08,"blended":0.76},"headline":{"value":0.76,"unit":"per 1M tokens","basis":"blended"},"endpoints":[{"id":"chat","method":"POST","path":"/chat/completions","name":"Chat Completions","status":"stable","recommended":true},{"id":"responses","method":"POST","path":"/responses","name":"Responses","status":"alpha"}],"family":"glm-4-6"}],"related":{"similar":[{"slug":"qwen-3-next-80b","name":"Qwen 3 Next 80b","provider":"alibaba","modality":"text","privacy":["private"],"created":1745903059,"headline":{"value":0.7375,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"llama-3-3-70b","name":"Llama 3.3 70B","provider":"meta","modality":"text","privacy":["private"],"created":1743897600,"headline":{"value":1.225,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"qwen-3-235b-a22b-thinking-2507","name":"Qwen 3 235B A22B Thinking 2507","provider":"alibaba","modality":"text","privacy":["private"],"created":1745903059,"headline":{"value":1.2125,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"qwen3-vl-235b","name":"Qwen3 VL 235B","provider":"alibaba","modality":"text","privacy":["private"],"created":1768521600,"headline":{"value":0.6325,"unit":"per 1M tokens","basis":"blended"},"variants":1}],"versions":[{"slug":"glm-5-3","name":"GLM 5.3","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1787011200,"headline":{"value":2.6875,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"glm-5-2","name":"GLM 5.2","provider":"zai","modality":"text","privacy":["private","e2ee"],"created":1781568000,"headline":{"value":2.15,"unit":"per 1M tokens","basis":"blended"},"variants":2},{"slug":"glm-5-1","name":"GLM 5.1","provider":"zai","modality":"text","privacy":["private"],"created":1775520000,"headline":{"value":2.365,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"glm-5","name":"GLM 5","provider":"zai","modality":"text","privacy":["private"],"created":1770768000,"headline":{"value":1.55,"unit":"per 1M tokens","basis":"blended"},"variants":1},{"slug":"glm-4-7","name":"GLM 4.7","provider":"zai","modality":"text","privacy":["private"],"created":1766534400,"headline":{"value":1.075,"unit":"per 1M tokens","basis":"blended"},"variants":1}]},"providers":{"zai":{"slug":"zai","name":"Z.ai","logo":"/images/icons/models/Zhipu.svg"},"alibaba":{"slug":"alibaba","name":"Alibaba Qwen","logo":"/images/icons/models/qwen.svg"},"meta":{"slug":"meta","name":"Meta","logo":"/images/icons/models/meta.svg"}},"faq":[{"q":"How much does the GLM 4.6 API cost?","a":"$0.43 per 1M input tokens and $1.75 per 1M output tokens, with cached input at $0.08 per 1M. Prices are in USD and can be paid in DIEM at parity."},{"q":"What is the GLM 4.6 model ID?","a":"Use `zai-org-glm-4.6` as the `model` parameter."},{"q":"Is the GLM 4.6 API private?","a":"GLM 4.6 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training."},{"q":"What is the context window of GLM 4.6?","a":"198K tokens of context, with up to 16K output tokens per response."},{"q":"What does GLM 4.6 support?","a":"GLM 4.6 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: low, medium and high (default low)."},{"q":"Which endpoint does the GLM 4.6 API use?","a":"Call `POST /chat/completions`. `/responses` (Alpha) is also supported."}]}} />

<div className="vx-static">
  <Accordion title="Plain-text specification">
    # GLM 4.6 API

    GLM 4.6 is a large language model by Z.ai, available on the Venice API as `zai-org-glm-4.6`. It runs privately, with zero data retention.

    GLM-4.6 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.

    ## GLM 4.6 API pricing

    | Model ID | Variant | Privacy | Price |
    | - | - | - | - |
    | `zai-org-glm-4.6` | Standard | Private | $0.43 input / $1.75 output per 1M tokens |

    ## GLM 4.6 specifications

    | Spec | Value |
    | - | - |
    | Provider | Z.ai |
    | Released | Apr 1, 2024 |
    | Privacy | Private |
    | Open weights | Yes |
    | Context window | 198K tokens |
    | Max output | 16K tokens |
    | Reasoning effort | low, medium, high |
    | Served precision | FP4 |

    ## How to use the GLM 4.6 API

    Send requests to `POST https://api.venice.ai/api/v1/chat/completions` with `"model": "zai-org-glm-4.6"` and your API key.

    ```bash theme={null}
    curl https://api.venice.ai/api/v1/chat/completions \
      -H "Authorization: Bearer $VENICE_API_KEY" \
      -H "Content-Type: application/json" \
      -d '{
        "model": "zai-org-glm-4.6",
        "messages": [{ "role": "user", "content": "Explain TEE attestation in two sentences." }],
        "reasoning_effort": "low"
      }'
    ```

    ## GLM 4.6 API FAQ

    ### How much does the GLM 4.6 API cost?

    $0.43 per 1M input tokens and $1.75 per 1M output tokens, with cached input at \$0.08 per 1M. Prices are in USD and can be paid in DIEM at parity.

    ### What is the GLM 4.6 model ID?

    Use `zai-org-glm-4.6` as the `model` parameter.

    ### Is the GLM 4.6 API private?

    GLM 4.6 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

    ### What is the context window of GLM 4.6?

    198K tokens of context, with up to 16K output tokens per response.

    ### What does GLM 4.6 support?

    GLM 4.6 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with `reasoning_effort`: low, medium and high (default low).

    ### Which endpoint does the GLM 4.6 API use?

    Call `POST /chat/completions`. `/responses` (Alpha) is also supported.

    ## Related models

    * [GLM 5.3 API](/models/glm-5-3): $1.75 input / $5.50 output per 1M tokens
    * [GLM 5.2 API](/models/glm-5-2): $1.40 input / $4.40 output per 1M tokens
    * [GLM 5.1 API](/models/glm-5-1): $1.54 input / $4.84 output per 1M tokens
    * [GLM 5 API](/models/glm-5): $1.00 input / $3.20 output per 1M tokens
    * [GLM 4.7 API](/models/glm-4-7): $0.55 input / $2.65 output per 1M tokens
    * [Qwen 3 Next 80b API](/models/qwen-3-next-80b): $0.35 input / $1.90 output per 1M tokens
    * [Llama 3.3 70B API](/models/llama-3-3-70b): $0.70 input / $2.80 output per 1M tokens
    * [Qwen 3 235B A22B Thinking 2507 API](/models/qwen-3-235b-a22b-thinking-2507): $0.45 input / $3.50 output per 1M tokens
    * [Qwen3 VL 235B API](/models/qwen3-vl-235b): $0.21 input / $1.90 output per 1M tokens
  </Accordion>
</div>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.