Skip to main content

MiniMax M2.5 API

MiniMax M2.5 is a large language model by MiniMax, available on the Venice API as minimax-m25. It runs privately, with zero data retention.MiniMax-M2.5 is a state-of-the-art large language model optimized for coding, agentic workflows, and modern application development with enhanced reasoning capabilities.

MiniMax M2.5 API pricing

MiniMax M2.5 specifications

How to use the MiniMax M2.5 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "minimax-m25" and your API key.

MiniMax M2.5 API FAQ

How much does the MiniMax M2.5 API cost?

0.27per1Minputtokensand0.27 per 1M input tokens and 0.95 per 1M output tokens, with cached input at $0.03 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the MiniMax M2.5 model ID?

Use minimax-m25 as the model parameter.

Is the MiniMax M2.5 API private?

MiniMax M2.5 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of MiniMax M2.5?

198K tokens of context, with up to 32K output tokens per response.

What does MiniMax M2.5 support?

MiniMax M2.5 supports function calling, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: low, medium and high (default low).

Which endpoint does the MiniMax M2.5 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models