Plain-text specification
Plain-text specification
MiMo-V2.5 API
MiMo-V2.5 is a large language model by Xiaomi, available on the Venice API asxiaomi-mimo-v2-5. It runs privately, with zero data retention.MiMo-V2.5 is Xiaomi’s native omnimodal model with strong agentic capabilities, supporting text, image, video, and audio understanding in a unified architecture. Built on a sparse Mixture-of-Experts backbone with 310B total and 15B active parameters, it delivers long-context reasoning up to 1M tokens, function calling, and multimodal perception.MiMo-V2.5 API pricing
MiMo-V2.5 specifications
How to use the MiMo-V2.5 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "xiaomi-mimo-v2-5" and your API key.MiMo-V2.5 API FAQ
How much does the MiMo-V2.5 API cost?
2.00 per 1M output tokens, with cached input at $0.08 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the MiMo-V2.5 model ID?
Usexiaomi-mimo-v2-5 as the model parameter.Is the MiMo-V2.5 API private?
MiMo-V2.5 is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of MiMo-V2.5?
1M tokens of context, with up to 64K output tokens per response.What does MiMo-V2.5 support?
MiMo-V2.5 supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the MiMo-V2.5 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Qwen 3.7 Plus API: 2.00 output per 1M tokens
- Venice Role Play Uncensored API: 2.00 output per 1M tokens
- MiniMax M2.7 API: 1.50 output per 1M tokens
- Gemini 3.5 Flash-Lite API: 3.13 output per 1M tokens