Skip to main content

Hermes 3 Llama 3.1 405b API

Hermes 3 Llama 3.1 405b is a large language model by Nous Research, available on the Venice API as hermes-3-llama-3.1-405b. It runs privately, with zero data retention.Hermes 3 405B is a frontier level, full parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user.

Hermes 3 Llama 3.1 405b API pricing

Hermes 3 Llama 3.1 405b specifications

How to use the Hermes 3 Llama 3.1 405b API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "hermes-3-llama-3.1-405b" and your API key.

Hermes 3 Llama 3.1 405b API FAQ

How much does the Hermes 3 Llama 3.1 405b API cost?

1.10per1Minputtokensand1.10 per 1M input tokens and 3.00 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.

What is the Hermes 3 Llama 3.1 405b model ID?

Use hermes-3-llama-3.1-405b as the model parameter.

Is the Hermes 3 Llama 3.1 405b API private?

Hermes 3 Llama 3.1 405b is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Hermes 3 Llama 3.1 405b?

125K tokens of context, with up to 16K output tokens per response.

What does Hermes 3 Llama 3.1 405b support?

Hermes 3 Llama 3.1 405b supports web search.

Which endpoint does the Hermes 3 Llama 3.1 405b API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models