Skip to main content

MiMo-V2.5 API

MiMo-V2.5 is a large language model by Xiaomi, available on the Venice API as xiaomi-mimo-v2-5. It runs privately, with zero data retention.MiMo-V2.5 is Xiaomi’s native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding in a unified architecture. Built on a sparse Mixture-of-Experts backbone with 310B total and 15B active parameters, it delivers long-context reasoning up to 1M tokens, function calling, and multimodal perception.

MiMo-V2.5 API pricing

MiMo-V2.5 specifications

How to use the MiMo-V2.5 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "xiaomi-mimo-v2-5" and your API key.

MiMo-V2.5 API FAQ

How much does the MiMo-V2.5 API cost?

0.40per1Minputtokensand0.40 per 1M input tokens and 2.00 per 1M output tokens, with cached input at $0.08 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the MiMo-V2.5 model ID?

Use xiaomi-mimo-v2-5 as the model parameter.

Is the MiMo-V2.5 API private?

MiMo-V2.5 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of MiMo-V2.5?

1M tokens of context, with up to 64K output tokens per response.

What does MiMo-V2.5 support?

MiMo-V2.5 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the MiMo-V2.5 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models