Skip to main content

Qwen 3.5 9B API

Qwen 3.5 9B is a large language model by Alibaba Qwen, available on the Venice API as qwen3-5-9b. It runs privately, with zero data retention.A 9B dense model with 262K native context window (extendable to 1M). Features Gated DeltaNet hybrid attention architecture for efficient long-context processing. Supports 201 languages, thinking/reasoning mode, and function calling.

Qwen 3.5 9B API pricing

Qwen 3.5 9B specifications

How to use the Qwen 3.5 9B API

Send requests to POST https://api.venice.ai/api/v1/chat/completions with "model": "qwen3-5-9b" and your API key.

Qwen 3.5 9B API FAQ

How much does the Qwen 3.5 9B API cost?

0.10per1Minputtokensand0.10 per 1M input tokens and 0.15 per 1M output tokens. Prices are in USD and can be paid in DIEM at parity.

What is the Qwen 3.5 9B model ID?

Use qwen3-5-9b as the model parameter.

Is the Qwen 3.5 9B API private?

Qwen 3.5 9B is private: requests run on infrastructure Venice controls with zero data retention, and prompts and outputs are never stored or used for training.

What is the context window of Qwen 3.5 9B?

250K tokens of context, with up to 32K output tokens per response.

What does Qwen 3.5 9B support?

Qwen 3.5 9B supports function calling, structured outputs, reasoning, image input and web search. Reasoning effort is adjustable with reasoning_effort: none, low, medium and high (default low).

Which endpoint does the Qwen 3.5 9B API use?

Call POST /chat/completions. /responses (Alpha) is also supported.

Related models