Plain-text specification
Plain-text specification
GLM 4.7 Flash API
GLM 4.7 Flash is a large language model by Z.ai, available on the Venice API aszai-org-glm-4.7-flash. It runs privately, with zero data retention.GLM-4.7-Flash is a fast inference variant of GLM-4.7, optimized for speed while maintaining strong reasoning capabilities. Ideal for applications requiring quick responses with good quality.GLM 4.7 Flash API pricing
GLM 4.7 Flash specifications
How to use the GLM 4.7 Flash API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "zai-org-glm-4.7-flash" and your API key.GLM 4.7 Flash API FAQ
How much does the GLM 4.7 Flash API cost?
0.40 per 1M output tokens, with cached input at $0.01 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GLM 4.7 Flash model ID?
Usezai-org-glm-4.7-flash as the model parameter.Is the GLM 4.7 Flash API private?
GLM 4.7 Flash is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of GLM 4.7 Flash?
125K tokens of context, with up to 16K output tokens per response.What does GLM 4.7 Flash support?
GLM 4.7 Flash supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the GLM 4.7 Flash API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- GLM 5.3 Flash API: 0.50 output per 1M tokens
- GLM 4.7 Flash Heretic API: 0.40 output per 1M tokens
- NVIDIA Nemotron 3 Nano 30B API: 0.30 output per 1M tokens
- Mistral Small 3.2 24B Instruct API: 0.25 output per 1M tokens
- Google Gemma 3 27B Instruct API: 0.20 output per 1M tokens