Plain-text specification
Plain-text specification
Llama 3.3 70B API
Llama 3.3 70B is a large language model by Meta, available on the Venice API asllama-3.3-70b. It runs privately, with zero data retention.Llama 3.3 70B API pricing
Llama 3.3 70B specifications
How to use the Llama 3.3 70B API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "llama-3.3-70b" and your API key.Llama 3.3 70B API FAQ
How much does the Llama 3.3 70B API cost?
2.80 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.What is the Llama 3.3 70B model ID?
Usellama-3.3-70b as the model parameter.Is the Llama 3.3 70B API private?
Llama 3.3 70B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Llama 3.3 70B?
125K tokens of context, with up to 4K output tokens per response.What does Llama 3.3 70B support?
Llama 3.3 70B supports function calling and web search.Which endpoint does the Llama 3.3 70B API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Llama 3.2 3B API: 0.60 output per 1M tokens
- Qwen 3 235B A22B Thinking 2507 API: 3.50 output per 1M tokens
- Kimi K2.5 API: 3.50 output per 1M tokens
- Hermes 3 Llama 3.1 405b API: 3.00 output per 1M tokens
- GLM 4.7 API: 2.65 output per 1M tokens