Plain-text specification
Plain-text specification
Qwen 3.6 35B A3B API
Qwen 3.6 35B A3B is a large language model by Alibaba Qwen, available on the Venice API asqwen3-6-35b-a3b. Private and end-to-end encrypted variants are available.Qwen 3.6 35B A3B is a fast mixture-of-experts model with 35B total parameters and ~3B active per token. Strong at agentic coding, STEM reasoning, and tool use, with a native 256K context window.Qwen 3.6 35B A3B API pricing
Qwen 3.6 35B A3B specifications
How to use the Qwen 3.6 35B A3B API
Send requests toPOST https://api.venice.ai/api/v1/chat/completions with "model": "qwen3-6-35b-a3b" and your API key.Qwen 3.6 35B A3B API FAQ
How much does the Qwen 3.6 35B A3B API cost?
1.00 per 1M output tokens. The E2EE variant costs 1.18 output. Prices are in USD and can be paid in DIEM at parity.What is the Qwen 3.6 35B A3B model ID?
Useqwen3-6-35b-a3b as the model parameter. Other variants: e2ee-qwen3-6-35b-a3b (E2EE).Is the Qwen 3.6 35B A3B API private?
The Standard variant is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training. The E2EE variant is end-to-end encrypted: prompts are encrypted on your device and decrypted only inside an attested hardware enclave, so neither Venice nor the GPU provider can read them.What is the context window of Qwen 3.6 35B A3B?
250K tokens of context, with up to 64K output tokens per response.What does Qwen 3.6 35B A3B support?
Qwen 3.6 35B A3B supports function calling, structured outputs, reasoning, image input and web search.Which endpoint does the Qwen 3.6 35B A3B API use?
CallPOST /chat/completions. /responses (Alpha) is also supported.Related models
- Qwen 3.5 35B A3B API: 1.25 output per 1M tokens
- Gemma 4 26B A4B Uncensored API: 0.88 output per 1M tokens
- GPT-4o Mini API: 0.75 output per 1M tokens
- Venice Uncensored 1.2 API: 0.90 output per 1M tokens
- Gemma 4 Uncensored API: 0.50 output per 1M tokens