Skip to main content

GLM 4.7 Flash Heretic API

GLM 4.7 Flash Heretic is a large language model by Community, available on the Venice API as olafangensan-glm-4.7-flash-heretic. It runs privately, with zero data retention.GLM-4.7-Flash-Heretic is an uncensored experimental variant of GLM-4.7-Flash, optimized for creative freedom and unfiltered dialogue with fast inference speed.

GLM 4.7 Flash Heretic API pricing

GLM 4.7 Flash Heretic specifications

How to use the GLM 4.7 Flash Heretic API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "olafangensan-glm-4.7-flash-heretic" and your API key.

GLM 4.7 Flash Heretic API FAQ

How much does the GLM 4.7 Flash Heretic API cost?

0.07per1Minputtokensand0.07 per 1M input tokens and 0.40 per 1M output tokens, with cached input at $0.035 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the GLM 4.7 Flash Heretic model ID?

Use olafangensan-glm-4.7-flash-heretic as the model parameter.

Is the GLM 4.7 Flash Heretic API private?

GLM 4.7 Flash Heretic is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of GLM 4.7 Flash Heretic?

200K tokens of context, with up to 24K output tokens per response.

What does GLM 4.7 Flash Heretic support?

GLM 4.7 Flash Heretic supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default low).

Which endpoint does the GLM 4.7 Flash Heretic API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models