Skip to main content

Mercury 2 API

Mercury 2 is a large language model by Inception, available on the Venice API as mercury-2. Requests are anonymized, so the provider never sees your identity.Mercury 2 is a diffusion-based reasoning LLM from Inception, delivering over 1,000 tokens per second — 5x faster than leading speed-optimized models — with strong reasoning, tool use, and structured output capabilities.

Mercury 2 API pricing

Mercury 2 specifications

How to use the Mercury 2 API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "mercury-2" and your API key.

Mercury 2 API FAQ

How much does the Mercury 2 API cost?

0.31per1Minputtokensand0.31 per 1M input tokens and 0.94 per 1M output tokens, with cached input at $0.031 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Mercury 2 model ID?

Use mercury-2 as the model parameter.

Is the Mercury 2 API private?

Mercury 2 is anonymized: Venice forwards requests to the provider without your identity, but the provider may retain prompt data, so use a private model for sensitive work.

What is the context window of Mercury 2?

125K tokens of context, with up to 50K output tokens per response.

What does Mercury 2 support?

Mercury 2 supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default high).

Which endpoint does the Mercury 2 API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models