Skip to main content

GPT-5.4 Mini API

GPT-5.4 Mini is a large language model by OpenAI, available on the Venice API as openai-gpt-54-mini. Requests are anonymized, so the provider never sees your identity.GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use.

GPT-5.4 Mini API pricing

GPT-5.4 Mini specifications

How to use the GPT-5.4 Mini API

Send requests to POST https://api.venice.ai/api/v1/responses with "model": "openai-gpt-54-mini" and your API key.

GPT-5.4 Mini API FAQ

How much does the GPT-5.4 Mini API cost?

0.94per1Minputtokensand0.94 per 1M input tokens and 5.63 per 1M output tokens, with cached input at $0.094 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the GPT-5.4 Mini model ID?

Use openai-gpt-54-mini as the model parameter.

Is the GPT-5.4 Mini API private?

GPT-5.4 Mini is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of GPT-5.4 Mini?

400K tokens of context, with up to 125K output tokens per response.

What does GPT-5.4 Mini support?

GPT-5.4 Mini supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium, high and xhigh (default high).

Which endpoint does the GPT-5.4 Mini API use?

Call POST /responses. /chat/completions is also supported. OpenAI reasoning models are designed around the Responses API: reasoning, tool calls and messages come back as typed output items. Venice’s /responses endpoint is in alpha.

Related models

  • DeepSeek V4 Pro API: 1.65input/1.65 input / 3.30 output per 1M tokens
  • GLM 5 Turbo API: 1.20input/1.20 input / 4.00 output per 1M tokens
  • GLM 5V Turbo API: 1.50input/1.50 input / 5.00 output per 1M tokens
  • GLM 5.1 API: 1.54input/1.54 input / 4.84 output per 1M tokens