Skip to main content

Qwen 3.5 35B A3B API

Qwen 3.5 35B A3B is a large language model by Alibaba Qwen, available on the Venice API as qwen3-5-35b-a3b. It runs privately, with zero data retention.Qwen 3.5 35B A3B is a highly efficient MoE model with 35B total parameters and only 3B active parameters. It surpasses the larger Qwen3-235B-A22B while being 6.7x smaller, excelling at reasoning, coding, and general knowledge tasks.

Qwen 3.5 35B A3B API pricing

Qwen 3.5 35B A3B specifications

How to use the Qwen 3.5 35B A3B API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "qwen3-5-35b-a3b" and your API key.

Qwen 3.5 35B A3B API FAQ

How much does the Qwen 3.5 35B A3B API cost?

0.31per1Minputtokensand0.31 per 1M input tokens and 1.25 per 1M output tokens, with cached input at $0.16 per 1M. Prices are in USD and can be paid in DIEM at parity.

What is the Qwen 3.5 35B A3B model ID?

Use qwen3-5-35b-a3b as the model parameter.

Is the Qwen 3.5 35B A3B API private?

Qwen 3.5 35B A3B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Qwen 3.5 35B A3B?

250K tokens of context, with up to 16K output tokens per response.

What does Qwen 3.5 35B A3B support?

Qwen 3.5 35B A3B supports function calling, structured outputs, reasoning, image input, web search and prompt caching. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default low).

Which endpoint does the Qwen 3.5 35B A3B API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models