Plain-text specification
Plain-text specification
Llama 3.2 3B API
Llama 3.2 3B is a large language model by Meta, available on the Venice API asllama-3.2-3b. It runs privately, with zero data retention.Llama 3.2 3B API pricing
Llama 3.2 3B specifications
How to use the Llama 3.2 3B API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "llama-3.2-3b" and your API key.Llama 3.2 3B API FAQ
How much does the Llama 3.2 3B API cost?
0.60 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.What is the Llama 3.2 3B model ID?
Usellama-3.2-3b as the model parameter.Is the Llama 3.2 3B API private?
Llama 3.2 3B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Llama 3.2 3B?
125K tokens of context, with up to 4K output tokens per response.What does Llama 3.2 3B support?
Llama 3.2 3B supports function calling and web search.Which endpoint does the Llama 3.2 3B API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Llama 3.3 70B API: 2.80 output per 1M tokens
- Qwen 3 235B A22B Instruct 2507 API: 0.75 output per 1M tokens
- Gemma 4 Uncensored API: 0.50 output per 1M tokens
- DeepSeek V3.2 API: 0.48 output per 1M tokens
- GPT-4o Mini API: 0.75 output per 1M tokens