Plain-text specification
Plain-text specification
NVIDIA Nemotron 3 Ultra API
NVIDIA Nemotron 3 Ultra is a large language model by NVIDIA, available on the Venice API asnvidia-nemotron-3-ultra-550b-a55b. It runs privately, with zero data retention.NVIDIA Nemotron 3 Ultra is built for frontier reasoning, orchestration, coding agents, deep research, and complex enterprise workflows. It delivers up to 5x faster inference and up to 30% lower cost for agentic workloads while supporting up to 1M token context.NVIDIA Nemotron 3 Ultra API pricing
NVIDIA Nemotron 3 Ultra specifications
How to use the NVIDIA Nemotron 3 Ultra API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "nvidia-nemotron-3-ultra-550b-a55b" and your API key.NVIDIA Nemotron 3 Ultra API FAQ
How much does the NVIDIA Nemotron 3 Ultra API cost?
3.13 per 1M output tokens, with cached input at $0.19 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the NVIDIA Nemotron 3 Ultra model ID?
Usenvidia-nemotron-3-ultra-550b-a55b as the model parameter.Is the NVIDIA Nemotron 3 Ultra API private?
NVIDIA Nemotron 3 Ultra is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of NVIDIA Nemotron 3 Ultra?
250K tokens of context, with up to 32K output tokens per response.What does NVIDIA Nemotron 3 Ultra support?
NVIDIA Nemotron 3 Ultra supports function calling, structured outputs, reasoning, web search and prompt caching. Reasoning effort is adjustable withreasoning_effort: none, low, medium and high (default low).Which endpoint does the NVIDIA Nemotron 3 Ultra API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Grok Build 0.1 API: 2.00 output per 1M tokens
- Seed 2.1 Turbo API: 3.13 output per 1M tokens
- Kimi K2.7 Code API: 3.50 output per 1M tokens
- Aion 3.0 Mini API: 1.75 output per 1M tokens