Skip to main content

MiMo-V2.6-Flash API

MiMo-V2.6-Flash is a large language model by Xiaomi, available on the Venice API as xiaomi-mimo-v2-6-flash. It runs privately, with zero data retention.MiMo-V2.6-Flash is Xiaomi’s efficiency-balanced omnimodal model: a sparse Mixture-of-Experts with 309B total and 15B activated parameters, natively understanding text, image, video, and audio with a 1M-token context window. Trained with mixed reinforcement learning across coding, general agents, visual, and cybersecurity tasks, it is built for long-horizon coding, tool use, and computer use.

MiMo-V2.6-Flash API pricing

MiMo-V2.6-Flash specifications

How to use the MiMo-V2.6-Flash API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "xiaomi-mimo-v2-6-flash" and your API key.

MiMo-V2.6-Flash API FAQ

How much does the MiMo-V2.6-Flash API cost?

0.17per1Minputtokensand0.17 per 1M input tokens and 0.35 per 1M output tokens, with cached input at $0.0037 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the MiMo-V2.6-Flash model ID?

Use xiaomi-mimo-v2-6-flash as the model parameter.

Is the MiMo-V2.6-Flash API private?

MiMo-V2.6-Flash is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of MiMo-V2.6-Flash?

1M tokens of context, with up to 128K output tokens per response.

What does MiMo-V2.6-Flash support?

MiMo-V2.6-Flash supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the MiMo-V2.6-Flash API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models