Plain-text specification
Plain-text specification
Mercury 2.5 API
Mercury 2.5 is a large language model by Inception, available on the Venice API asmercury-2-5. Requests are anonymized, so the provider never sees your identity.Mercury 2.5 is a diffusion-based reasoning model from Inception with fast parallel token generation, tunable reasoning, tool calling, and structured output support.Mercury 2.5 API pricing
Mercury 2.5 specifications
How to use the Mercury 2.5 API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "mercury-2-5" and your API key.Mercury 2.5 API FAQ
How much does the Mercury 2.5 API cost?
0.19 per 1M output tokens, with cached input at $0.005 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Mercury 2.5 model ID?
Usemercury-2-5 as the model parameter.Is the Mercury 2.5 API private?
Mercury 2.5 is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.What is the context window of Mercury 2.5?
260K tokens of context, with up to 64K output tokens per response.What does Mercury 2.5 support?
Mercury 2.5 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default high).Which endpoint does the Mercury 2.5 API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Mercury 2 API: 0.94 output per 1M tokens
- Qwen 2.5 7B API: 0.13 output per 1M tokens
- Qwen 3.5 9B API: 0.15 output per 1M tokens
- NVIDIA Nemotron 3 Nano 30B API: 0.30 output per 1M tokens
- Mistral Small 3.2 24B Instruct API: 0.25 output per 1M tokens