Skip to main content

Mercury 2.5 API

Mercury 2.5 is a large language model by Inception, available on the Venice API as mercury-2-5. Requests are anonymized, so the provider never sees your identity.Mercury 2.5 is a diffusion-based reasoning model from Inception with fast parallel token generation, tunable reasoning, tool calling, and structured output support.

Mercury 2.5 API pricing

Mercury 2.5 specifications

How to use the Mercury 2.5 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "mercury-2-5" and your API key.

Mercury 2.5 API FAQ

How much does the Mercury 2.5 API cost?

0.05per1Minputtokensand0.05 per 1M input tokens and 0.19 per 1M output tokens, with cached input at $0.005 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Mercury 2.5 model ID?

Use mercury-2-5 as the model parameter.

Is the Mercury 2.5 API private?

Mercury 2.5 is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of Mercury 2.5?

260K tokens of context, with up to 64K output tokens per response.

What does Mercury 2.5 support?

Mercury 2.5 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the Mercury 2.5 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models