Plain-text specification
Plain-text specification
Mercury 2 API
Mercury 2 is a large language model by Inception, available on the Venice API asmercury-2. Requests are anonymized, so the provider never sees your identity.Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured output capabilities.Mercury 2 API pricing
Mercury 2 specifications
How to use the Mercury 2 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "mercury-2" and your API key.Mercury 2 API FAQ
How much does the Mercury 2 API cost?
0.94 per 1M output tokens, with cached input at $0.031 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Mercury 2 model ID?
Usemercury-2 as the model parameter.Is the Mercury 2 API private?
Mercury 2 is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of Mercury 2?
125K tokens of context, with up to 50K output tokens per response.What does Mercury 2 support?
Mercury 2 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the Mercury 2 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Mercury 2.5 API: 0.19 output per 1M tokens
- MiniMax M2.5 API: 0.95 output per 1M tokens
- Qwen 3.5 35B A3B API: 1.25 output per 1M tokens
- Qwen3 VL 30B A3B API: 0.90 output per 1M tokens
- MiniMax M3 Preview API: 1.20 output per 1M tokens