Plain-text specification
Plain-text specification
GLM 5V Turbo API
GLM 5V Turbo is a large language model by Z.ai, available on the Venice API asz-ai-glm-5v-turbo. Requests are anonymized, so the provider never sees your identity.GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks with image, video, and text inputs.GLM 5V Turbo API pricing
GLM 5V Turbo specifications
How to use the GLM 5V Turbo API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "z-ai-glm-5v-turbo" and your API key.GLM 5V Turbo API FAQ
How much does the GLM 5V Turbo API cost?
5.00 per 1M output tokens, with cached input at $0.30 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GLM 5V Turbo model ID?
Usez-ai-glm-5v-turbo as the model parameter.Is the GLM 5V Turbo API private?
GLM 5V Turbo is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of GLM 5V Turbo?
200K tokens of context, with up to 32K output tokens per response.What does GLM 5V Turbo support?
GLM 5V Turbo supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the GLM 5V Turbo API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- GPT-5.4 Mini API: 5.63 output per 1M tokens
- DeepSeek V4 Pro API: 3.30 output per 1M tokens
- Inkling API: 5.06 output per 1M tokens
- DeepSeek V4 Pro 0813 API: 4.95 output per 1M tokens