Plain-text specification
Plain-text specification
NVIDIA Nemotron 3 Nano 30B API
NVIDIA Nemotron 3 Nano 30B is a large language model by NVIDIA, available on the Venice API asnvidia-nemotron-3-nano-30b-a3b. It runs privately, with zero data retention.NVIDIA Nemotron 3 Nano 30B is a compact and efficient language model from NVIDIA, optimized for fast inference while maintaining strong performance across diverse tasks.NVIDIA Nemotron 3 Nano 30B API pricing
NVIDIA Nemotron 3 Nano 30B specifications
How to use the NVIDIA Nemotron 3 Nano 30B API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "nvidia-nemotron-3-nano-30b-a3b" and your API key.NVIDIA Nemotron 3 Nano 30B API FAQ
How much does the NVIDIA Nemotron 3 Nano 30B API cost?
0.30 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.What is the NVIDIA Nemotron 3 Nano 30B model ID?
Usenvidia-nemotron-3-nano-30b-a3b as the model parameter.Is the NVIDIA Nemotron 3 Nano 30B API private?
NVIDIA Nemotron 3 Nano 30B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of NVIDIA Nemotron 3 Nano 30B?
125K tokens of context, with up to 16K output tokens per response.What does NVIDIA Nemotron 3 Nano 30B support?
NVIDIA Nemotron 3 Nano 30B supports function calling, structured outputs and web search.Which endpoint does the NVIDIA Nemotron 3 Nano 30B API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Mistral Small 3.2 24B Instruct API: 0.25 output per 1M tokens
- GLM 4.7 Flash API: 0.40 output per 1M tokens
- OpenAI GPT OSS 120B API: 0.30 output per 1M tokens
- GLM 4.7 Flash Heretic API: 0.40 output per 1M tokens