Skip to main content

Llama 3.3 70B API

Llama 3.3 70B is a large language model by Meta, available on the Venice API as llama-3.3-70b. It runs privately, with zero data retention.

Llama 3.3 70B API pricing

Llama 3.3 70B specifications

How to use the Llama 3.3 70B API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "llama-3.3-70b" and your API key.

Llama 3.3 70B API FAQ

How much does the Llama 3.3 70B API cost?

0.70per1Minputtokensand0.70 per 1M input tokens and 2.80 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.

What is the Llama 3.3 70B model ID?

Use llama-3.3-70b as the model parameter.

Is the Llama 3.3 70B API private?

Llama 3.3 70B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Llama 3.3 70B?

125K tokens of context, with up to 4K output tokens per response.

What does Llama 3.3 70B support?

Llama 3.3 70B supports function calling and web search.

Which endpoint does the Llama 3.3 70B API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models