Skip to main content

NVIDIA Nemotron 3 Nano 30B API

NVIDIA Nemotron 3 Nano 30B is a large language model by NVIDIA, available on the Venice API as nvidia-nemotron-3-nano-30b-a3b. It runs privately, with zero data retention.NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.

NVIDIA Nemotron 3 Nano 30B API pricing

NVIDIA Nemotron 3 Nano 30B specifications

How to use the NVIDIA Nemotron 3 Nano 30B API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "nvidia-nemotron-3-nano-30b-a3b" and your API key.

NVIDIA Nemotron 3 Nano 30B API FAQ

How much does the NVIDIA Nemotron 3 Nano 30B API cost?

0.075per1Minputtokensand0.075 per 1M input tokens and 0.30 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.

What is the NVIDIA Nemotron 3 Nano 30B model ID?

Use nvidia-nemotron-3-nano-30b-a3b as the model parameter.

Is the NVIDIA Nemotron 3 Nano 30B API private?

NVIDIA Nemotron 3 Nano 30B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of NVIDIA Nemotron 3 Nano 30B?

125K tokens of context, with up to 16K output tokens per response.

What does NVIDIA Nemotron 3 Nano 30B support?

NVIDIA Nemotron 3 Nano 30B supports function calling, structured outputs and web search.

Which endpoint does the NVIDIA Nemotron 3 Nano 30B API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models