Plain-text specification
Plain-text specification
Hermes 3 Llama 3.1 405b API
Hermes 3 Llama 3.1 405b is a large language model by Nous Research, available on the Venice API ashermes-3-llama-3.1-405b. It runs privately, with zero data retention.Hermes 3 405B is a frontier level, full parameter finetune of the Llama-3.1 405B foundation model, focused on aligning LLMs to the user, with powerful steering capabilities and control given to the end user.Hermes 3 Llama 3.1 405b API pricing
Hermes 3 Llama 3.1 405b specifications
How to use the Hermes 3 Llama 3.1 405b API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "hermes-3-llama-3.1-405b" and your API key.Hermes 3 Llama 3.1 405b API FAQ
How much does the Hermes 3 Llama 3.1 405b API cost?
3.00 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.What is the Hermes 3 Llama 3.1 405b model ID?
Usehermes-3-llama-3.1-405b as the model parameter.Is the Hermes 3 Llama 3.1 405b API private?
Hermes 3 Llama 3.1 405b is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Hermes 3 Llama 3.1 405b?
125K tokens of context, with up to 16K output tokens per response.What does Hermes 3 Llama 3.1 405b support?
Hermes 3 Llama 3.1 405b supports web search.Which endpoint does the Hermes 3 Llama 3.1 405b API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Gemini 3 Flash Preview API: 3.75 output per 1M tokens
- GLM 5 API: 3.20 output per 1M tokens
- Qwen 3.5 397B API: 4.50 output per 1M tokens
- Grok 4.20 API: 2.83 output per 1M tokens