Plain-text specification
Plain-text specification
MiMo-V2.6-Flash API
MiMo-V2.6-Flash is a large language model by Xiaomi, available on the Venice API asxiaomi-mimo-v2-6-flash. It runs privately, with zero data retention.MiMo-V2.6-Flash is Xiaomi’s efficiency-balanced omnimodal model: a sparse Mixture-of-Experts with 309B total and 15B activated parameters, natively understanding text, image, video, and audio with a 1M-token context window. Trained with mixed reinforcement learning across coding, general agents, visual, and cybersecurity tasks, it is built for long-horizon coding, tool use, and computer use.MiMo-V2.6-Flash API pricing
MiMo-V2.6-Flash specifications
How to use the MiMo-V2.6-Flash API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "xiaomi-mimo-v2-6-flash" and your API key.MiMo-V2.6-Flash API FAQ
How much does the MiMo-V2.6-Flash API cost?
0.35 per 1M output tokens, with cached input at $0.0037 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the MiMo-V2.6-Flash model ID?
Usexiaomi-mimo-v2-6-flash as the model parameter.Is the MiMo-V2.6-Flash API private?
MiMo-V2.6-Flash is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of MiMo-V2.6-Flash?
1M tokens of context, with up to 128K output tokens per response.What does MiMo-V2.6-Flash support?
MiMo-V2.6-Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the MiMo-V2.6-Flash API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Qwen 3.8 Flash API: 0.49 output per 1M tokens
- DeepSeek V4 Flash 0731 API: 0.35 output per 1M tokens
- GLM 5.3 Flash API: 0.50 output per 1M tokens
- GPT-6 Luna API: 0.63 output per 1M tokens