Skip to main content

NVIDIA Nemotron 3 Ultra API

NVIDIA Nemotron 3 Ultra is a large language model by NVIDIA, available on the Venice API as nvidia-nemotron-3-ultra-550b-a55b. It runs privately, with zero data retention.NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.

NVIDIA Nemotron 3 Ultra API pricing

NVIDIA Nemotron 3 Ultra specifications

How to use the NVIDIA Nemotron 3 Ultra API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "nvidia-nemotron-3-ultra-550b-a55b" and your API key.

NVIDIA Nemotron 3 Ultra API FAQ

How much does the NVIDIA Nemotron 3 Ultra API cost?

0.63per1Minputtokensand0.63 per 1M input tokens and 3.13 per 1M output tokens, with cached input at $0.19 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the NVIDIA Nemotron 3 Ultra model ID?

Use nvidia-nemotron-3-ultra-550b-a55b as the model parameter.

Is the NVIDIA Nemotron 3 Ultra API private?

NVIDIA Nemotron 3 Ultra is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of NVIDIA Nemotron 3 Ultra?

250K tokens of context, with up to 32K output tokens per response.

What does NVIDIA Nemotron 3 Ultra support?

NVIDIA Nemotron 3 Ultra supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default low).

Which endpoint does the NVIDIA Nemotron 3 Ultra API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models