Plain-text specification
Plain-text specification
GLM 5 Turbo API
GLM 5 Turbo is a large language model by Z.ai, available on the Venice API asz-ai-glm-5-turbo. Requests are anonymized, so the provider never sees your identity.GLM-5 Turbo is a fast inference model from Z.ai tuned for strong performance in agent-driven environments and production coding workflows.GLM 5 Turbo API pricing
GLM 5 Turbo specifications
How to use the GLM 5 Turbo API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "z-ai-glm-5-turbo" and your API key.GLM 5 Turbo API FAQ
How much does the GLM 5 Turbo API cost?
4.00 per 1M output tokens, with cached input at $0.24 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GLM 5 Turbo model ID?
Usez-ai-glm-5-turbo as the model parameter.Is the GLM 5 Turbo API private?
GLM 5 Turbo is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of GLM 5 Turbo?
200K tokens of context, with up to 32K output tokens per response.What does GLM 5 Turbo support?
GLM 5 Turbo supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the GLM 5 Turbo API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Grok 4.20 API: 2.83 output per 1M tokens
- Grok 4.20 Multi-Agent API: 2.83 output per 1M tokens
- Grok 4.3 API: 2.83 output per 1M tokens
- GPT-5.4 Mini API: 5.63 output per 1M tokens