Plain-text specification
Plain-text specification
GPT-4o Mini API
GPT-4o Mini is a large language model by OpenAI, available on the Venice API asopenai-gpt-4o-mini-2024-07-18. Requests are anonymized, so the provider never sees your identity.OpenAI’s cost-efficient small model that delivers GPT-4 level intelligence at a fraction of the cost. Ideal for high-volume applications requiring strong reasoning. Version: 2024-07-18.GPT-4o Mini API pricing
GPT-4o Mini specifications
How to use the GPT-4o Mini API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "openai-gpt-4o-mini-2024-07-18" and your API key.GPT-4o Mini API FAQ
How much does the GPT-4o Mini API cost?
0.75 per 1M output tokens, with cached input at $0.094 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the GPT-4o Mini model ID?
Useopenai-gpt-4o-mini-2024-07-18 as the model parameter.Is the GPT-4o Mini API private?
GPT-4o Mini is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of GPT-4o Mini?
125K tokens of context, with up to 16K output tokens per response.What does GPT-4o Mini support?
GPT-4o Mini supports function calling, structured outputs, image input, web search and prompt caching.Which endpoint does the GPT-4o Mini API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Qwen 3.6 35B A3B API: 1.00 output per 1M tokens
- Venice Uncensored 1.2 API: 0.90 output per 1M tokens
- Gemma 4 26B A4B Uncensored API: 0.88 output per 1M tokens
- DeepSeek V3.2 API: 0.48 output per 1M tokens