Plain-text specification
Plain-text specification
GLM 4.7 API
GLM 4.7 is a large language model by Z.ai, available on the Venice API aszai-org-glm-4.7. It runs privately, with zero data retention.GLM-4.7 is a large language model developed by Zhiyuan AI, featuring strong reasoning capabilities and support for multiple languages. Supports the largest context window for processing extensive text and detailed analysis.GLM 4.7 API pricing
GLM 4.7 specifications
How to use the GLM 4.7 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "zai-org-glm-4.7" and your API key.GLM 4.7 API FAQ
How much does the GLM 4.7 API cost?
2.65 per 1M output tokens, with cached input at $0.11 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GLM 4.7 model ID?
Usezai-org-glm-4.7 as the model parameter.Is the GLM 4.7 API private?
GLM 4.7 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of GLM 4.7?
198K tokens of context, with up to 16K output tokens per response.What does GLM 4.7 support?
GLM 4.7 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: low, medium and high (default low).Which endpoint does the GLM 4.7 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- GLM 5.3 API: 5.50 output per 1M tokens
- GLM 5.2 API: 4.40 output per 1M tokens
- GLM 5.1 API: 4.84 output per 1M tokens
- GLM 5 API: 3.20 output per 1M tokens
- GLM 4.6 API: 1.75 output per 1M tokens
- Qwen 3.6 27B API: 3.25 output per 1M tokens
- Kimi K2.5 API: 3.50 output per 1M tokens
- Gemini 3.5 Flash-Lite API: 3.13 output per 1M tokens
- Venice Role Play Uncensored API: 2.00 output per 1M tokens