Plain-text specification
Plain-text specification
Qwen 3.8 2.4T API
Qwen 3.8 2.4T is a large language model by Alibaba Qwen, available on the Venice API asqwen-3-8-2-4t-a95b. It runs privately, with zero data retention.Qwen 3.8 2.4T is Alibaba’s open-weight 2.4-trillion-parameter MoE model (95B active), with major gains in software engineering, research, and long-horizon agentic tasks. It is text-only, requires thinking mode, and supports a 262K-token context window.Qwen 3.8 2.4T API pricing
Qwen 3.8 2.4T specifications
How to use the Qwen 3.8 2.4T API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "qwen-3-8-2-4t-a95b" and your API key.Qwen 3.8 2.4T API FAQ
How much does the Qwen 3.8 2.4T API cost?
7.50 per 1M output tokens, with cached input at $0.31 per 1M. Prices are in USD and can be paid in DIEM at parity.What is the Qwen 3.8 2.4T model ID?
Useqwen-3-8-2-4t-a95b as the model parameter.Is the Qwen 3.8 2.4T API private?
Qwen 3.8 2.4T is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.What is the context window of Qwen 3.8 2.4T?
256K tokens of context, with up to 64K output tokens per response.What does Qwen 3.8 2.4T support?
Qwen 3.8 2.4T supports function calling, reasoning, web search and prompt caching.Which endpoint does the Qwen 3.8 2.4T API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Grok 4.6 API: 6.80 output per 1M tokens
- Grok 4.7 API: 6.80 output per 1M tokens
- Grok 4.5 API: 6.80 output per 1M tokens
- Gemini 3.5 Flash API: 9.45 output per 1M tokens